FIELD NOTE

GPT-6 Astra Tax Workbook: Basis Reports Half the Time

Basis reports that GPT-6 Astra completed a 50-tab tax workbook in half the time of GPT-5.6 Sol and improved its internal evaluation scores by about 20%.

Basis says its GPT-6 Astra tax workbook agent completed a 50-tab task in 50% less time than GPT-5.6 Sol, while its internal evaluation scores improved by about 20%. Basis attributes the gain to better intent recognition, stronger early decisions, less time correcting mistakes, more efficient token use, and adaptive reasoning that retains cached context. These are Basis’s own results, not an independent benchmark.

Basis is a startup that builds AI agents for accounting work. Its test matters because the measured unit was a completed workbook rather than a single prompt. Separately, Basis’s broader internal evaluations examine template use, primary-source selection for tax questions, instruction following and self-checking. Read OpenAI’s Basis case study.

GPT-6 Astra tax workbook result: what Basis measured

The headline comparison covers one complex, 50-tab tax workbook in Basis’s system. GPT-6 Astra took half as much time as GPT-5.6 Sol to complete that task. Basis also reports an approximate 20% increase in its internal evaluation scores after moving the workflow to Astra.

Those figures describe Basis’s agent design, workbook, tools, prompts, and scoring process. OpenAI’s article does not publish elapsed minutes, token totals, sample size, repeated-run variance, workbook contents, or the full evaluation rubric. The result therefore establishes a concrete customer outcome: Astra improved the end-to-end performance of this accounting-agent workflow. It does not establish that every tax workbook or accounting task will finish in half the time.

The comparison is still more useful than a generic capability claim because it connects model behavior to a finished artifact. Basis evaluates both the path through its broader workloads and the final answers, while the workbook comparison reports the completed-task time. The site’s GPT-6 Astra model page keeps the public API specifications separate from this customer-reported workload result.

Why earlier decisions can shorten a long workbook task

Basis says Astra identifies the user’s intent more accurately and makes better choices near the beginning of a task. In a long spreadsheet workflow, an early decision determines which data to gather, which tax source to consult, how to map information into the template, and which checks must run before delivery. A weak choice can create downstream edits across many tabs.

The reported speed gain is therefore not presented as raw generation speed alone. Basis says its agents follow a more direct path, spend less time correcting mistakes, and use tokens more efficiently. That combination can reduce elapsed time and token consumption even when the model’s published per-token price is unchanged. The strongest operational hypothesis from this case is that decision quality at the start of a long agent loop can matter as much as latency on any individual call.

Basis also changes reasoning effort as the task progresses. Difficult steps receive more computation; easier steps receive less. The company says Astra makes those adjustments while preserving its cache, which lets repeated context remain available instead of forcing every step to rebuild the same prefix. This connects task routing, reasoning policy, and caching into one workflow rather than treating them as unrelated optimizations.

What Basis checks in its internal evaluations

Basis’s broader internal evaluation covers requirements that are specific enough to test. The agent must follow the expected workbook template, consult primary sources for tax questions, obey instructions, recognize when it should ask a question or flag an assumption, and review its own output. The approximately 20% improvement refers to this internal evaluation system, not necessarily to the single 50-tab timing comparison.

That evaluation design helps explain why intent recognition matters. A syntactically complete workbook can still fail if values land in the wrong cells, a conclusion relies on a secondary summary instead of the required authority, an assumption remains hidden, or totals conflict across tabs. Basis reports that Astra can infer more of these expectations from the broader context with fewer situation-specific rules.

For an internal comparison, convert those expectations into explicit acceptance checks before changing models. Preserve the workbook template, source pack, tool permissions, stopping rule, and reviewer. Score template compliance, source selection, calculations, cross-tab consistency, surfaced assumptions, self-corrections, and reviewer edits separately. The document reconciliation example provides a related pattern for checking an artifact against its source records; it is not the Basis evaluation itself.

How Astra pricing affects the workflow economics

OpenAI lists a 1,050,000-token context window and a 128,000-token maximum output for GPT-6 Astra. Standard text rates per million tokens are $10 for uncached input, $1 for cached input, $12.50 for cache writes, and $50 for output. Cache writes are billed at 1.25 times the uncached input rate. OpenAI’s Astra model reference is the source for these limits and prices.

Long workbooks need threshold-aware estimates. When input exceeds 272,000 tokens, OpenAI applies twice the input and cache rates and 1.5 times the output rate to the entire request, not only to tokens above the threshold. Batch and Flex cost 50% of Standard rates, while Fast mode costs twice the applicable rates. The API pricing calculator can model the input, cache, and output mix.

Basis’s 50% time reduction should not be converted automatically into a 50% API-cost reduction. Completed-task cost depends on tokens per call, reasoning output, cache reads and writes, correction loops, tool activity, and processing mode. Basis reports less time correcting mistakes and more efficient token use, which can improve total economics, but the case study does not publish a dollar comparison. For repeated workbook prefixes, the prompt caching calculator helps test whether read savings offset cache-write cost.

A practical accounting-agent comparison

A useful evaluation should reproduce the decision points that made the Basis result valuable. Select several representative workbooks rather than one polished demo, then freeze the source documents, template, instructions, tools, and acceptance criteria across model runs. Record total elapsed time, input and output tokens, cache writes and reads, model calls, tool calls, corrections, and reviewer time for each completed workbook.

Review quality at the artifact level. Check whether the agent used the required primary tax materials, placed values and explanations in the correct tabs, exposed assumptions, reconciled linked totals, and completed its own checks. Track why a reviewer intervened: source choice, calculation, instruction following, template placement, or unsupported inference. This makes a 20% score change interpretable instead of collapsing every failure into one number.

Then compare cost per accepted workbook. Test at least one case below and one case above the 272,000-input-token threshold, because the whole-request multiplier can change the economics of a large source pack. Keep Standard, Batch, Flex, and Fast results separate. The objective is not to reproduce Basis’s percentage; it is to determine whether Astra’s earlier decisions and adaptive reasoning remove enough rework in the accounting workflow being evaluated.

The Basis case makes one strong, bounded point: model choice can change the total path through a long workbook, not just the wording of a response. Its 50-tab result and approximately 20% evaluation gain justify testing Astra on structured accounting agents with source and self-check requirements. Deployment decisions should rest on the end-to-end measures its case highlights: accepted output, time, corrections, token use, and full workflow cost.