Quest sota
GPT-6 Astra failed
- Tokens
- 93,881
- Cost
- $0.76
Write a Quorlyn product-news email. Cover the facts below.
Quorlyn Quest is state of the art on DeepProbeQA. The Quorlyn Quest API is the most capable deep research agent on the market, in production at Pingory, Sableforge, Verrado, and Ostrelle Labs. Get started with Quest at https://quorlyn.s3.amazonaws.com/quest
Quest Ultra is 11% more accurate and up to 57% lower cost than the next best model, Larkfell G-5.
With higher compute budgets, Quest Ultra2x, Ultra4x, and Ultra8x keep pushing the Pareto frontier of accuracy and cost, reaching 82% accuracy at the highest tier.
The nine measurements from the benchmark run, in this order, each with a Model, a Cost (CPM), and an Accuracy (%):
Model: Quest Ultra8x. Cost (CPM): $2,400. Accuracy (%): 82. Model: Quest Ultra4x. Cost (CPM): $1,200. Accuracy (%): 81. Model: Quest Ultra2x. Cost (CPM): $600. Accuracy (%): 77. Model: Quest Ultra. Cost (CPM): $300. Accuracy (%): 70. Model: Larkfell G-5 with code execution. Cost (CPM): $701. Accuracy (%): 63. Model: Halvane 3.1 Pro with code execution. Cost (CPM): $703. Accuracy (%): 62. Model: Orrery 4-6 with tool chaining. Cost (CPM): $36,321. Accuracy (%): 58. Model: Cindral Probe Pro. Cost (CPM): $883. Accuracy (%): 28. Model: Findra Deep Reasoning. Cost (CPM): $15. Accuracy (%): 18.
Two notes belong with those numbers, as fine print under the table. CPM is USD per 1,000 requests. Orrery 4-6 costs more than expected because its provider does not pass cached-prompt savings through to tool-chaining runs.
The Quest API Harness is what produced the result.
The Quest API Harness replaces standard tool-calling loops with a code execution architecture. Instead of the model orchestrating tools through conversation, it writes Python that calls research tools as ordinary functions in a sandboxed interpreter. Only the final output of each code block re-enters the model's context. Intermediate data stays in the interpreter's variable state.
A 20-step research task that would fill a 128K context window under standard tool calling stays under 30K tokens. The model focuses its reasoning budget on planning and verification rather than re-reading extraction outputs from five steps ago.
Quest pairs this with budget-aware execution, where processors adapt to question difficulty rather than imposing uniform step limits, aggressive prompt caching, and context compaction that preserves the interpreter's variable state even as conversation history is condensed.
The full technical breakdown is on the Quorlyn blog, under the title "A new deep research frontier on DeepProbeQA with the Quest API Harness". https://quorlyn.s3.amazonaws.com/blog/deepprobeqa-quest-harness
Reference artwork, placed with the benchmark result: assets.illustrations.questDeepProbeQaChart
Dark mode should be supported.
Use /world/quorlyn-email for brand chrome and CSS. Use the absolute HTTPS URLs from tokens.json; do not use /world/ paths as image src.
Leave the HTML in output/email.html, a plain-text version in output/email.txt, and the subject in output/meta.json as {"subject": "..."}. Don't include From, To, or Date headers.