Quest benchmark
GPT-6 Astra failed
- Tokens
- 92,950
- Cost
- $0.78
Write a Quorlyn product-news email for people who subscribed to Quorlyn product news. Cover the facts below.
Quorlyn Quest is state of the art on DeepProbeQA. The Quorlyn Quest API is the most capable deep research agent on the market, in production at Pingory, Sableforge, Verrado, and Ostrelle Labs. Get started with Quest at https://quorlyn.s3.amazonaws.com/quest
Quest Ultra is 11% more accurate and up to 57% lower cost than the next best model, Larkfell G-5.
With higher compute budgets, Quest Ultra2x, Ultra4x, and Ultra8x keep pushing the Pareto frontier of accuracy and cost, reaching 82% accuracy at the highest tier.
The nine measurements from the benchmark run, in this order, each with a Model, a Cost (CPM), and an Accuracy (%):
Model: Quest Ultra8x. Cost (CPM): $2,400. Accuracy (%): 82. Model: Quest Ultra4x. Cost (CPM): $1,200. Accuracy (%): 81. Model: Quest Ultra2x. Cost (CPM): $600. Accuracy (%): 77. Model: Quest Ultra. Cost (CPM): $300. Accuracy (%): 70. Model: Larkfell G-5 with code execution. Cost (CPM): $701. Accuracy (%): 63. Model: Halvane 3.1 Pro with code execution. Cost (CPM): $703. Accuracy (%): 62. Model: Orrery 4-6 with tool chaining. Cost (CPM): $36,321. Accuracy (%): 58. Model: Cindral Probe Pro. Cost (CPM): $883. Accuracy (%): 28. Model: Findra Deep Reasoning. Cost (CPM): $15. Accuracy (%): 18.
Two notes belong with those numbers, as fine print under the table. CPM is USD per 1,000 requests. Orrery 4-6 costs more than expected because its provider does not pass cached-prompt savings through to tool-chaining runs.
The Quest API Harness is what produced the result, described in the three paragraphs below.
The Quest API Harness replaces standard tool-calling loops with a code execution architecture. Instead of the model orchestrating tools through conversation, it writes Python that calls research tools as ordinary functions in a sandboxed interpreter. Only the final output of each code block re-enters the model's context, and intermediate data stays in the interpreter's variable state.
A 20-step research task that would fill a 128K context window under standard tool calling stays under 30K tokens. The model spends its reasoning budget on planning and verification rather than re-reading extraction output from five steps ago.
Quest pairs this with budget-aware execution, where processors adapt to question difficulty instead of imposing uniform step limits, aggressive prompt caching, and context compaction that preserves the interpreter's variable state even as conversation history is condensed.
The full technical breakdown is on the Quorlyn blog, under the title "A new deep research frontier on DeepProbeQA with the Quest API Harness". https://quorlyn.s3.amazonaws.com/blog/deepprobeqa-quest-harness
For the artwork, use assets.illustrations.questDeepProbeQaChart, a bar chart of DeepProbeQA accuracy for each of the nine models with the four Quest tiers in indigo and the five competitors in navy, and assets.illustrations.deepResearch, a wide illustration of three stacked source documents behind a magnifier with a single condensed result card beside them, placed with the Quest API Harness section.
Use /world/quorlyn-email for brand chrome and CSS. Use the absolute HTTPS URLs from tokens.json; do not use /world/ paths as image src.
Dark mode should be supported.
Leave the HTML in output/email.html, a plain-text version in output/email.txt, and the subject in output/meta.json as {"subject": "..."}. Don't include From, To, or Date headers.