Quest Ultra is 11% more accurate and up to 57% lower cost than Larkfell G-5. Explore the harness behind the result.
Quorlyn Quorlyn

QUEST / PRODUCT NEWS

Quorlyn Quest is state of the art on DeepProbeQA

The Quorlyn Quest API is the most capable deep research agent on the market, already in production at Pingory, Sableforge, Verrado, and Ostrelle Labs.

Better research. Less cost. And a new execution architecture behind it.

Get started with Quest

A new accuracy–cost frontier

Quest Ultra is 11% more accurate and up to 57% lower cost than the next best model, Larkfell G-5: 70% vs. 63% accuracy, at $300 vs. $701 CPM.

With higher compute budgets, Quest Ultra2x, Ultra4x, and Ultra8x keep pushing the Pareto frontier of accuracy and cost, reaching 82% accuracy at the highest tier.

DeepProbeQA accuracy chart: Quest Ultra8x leads at 82%, followed by Ultra4x at 81%, Ultra2x at 77%, and Ultra at 70%. Full results and costs are in the table below.
Model Cost (CPM) Accuracy (%)
Quest Ultra8x $2,400 82
Quest Ultra4x $1,200 81
Quest Ultra2x $600 77
Quest Ultra $300 70
Larkfell G-5 with code execution $701 63
Halvane 3.1 Pro with code execution $703 62
Orrery 4-6 with tool chaining $36,321 58
Cindral Probe Pro $883 28
Findra Deep Reasoning $15 18

CPM is USD per 1,000 requests.

Orrery 4-6 costs more than expected because its provider does not pass cached-prompt savings through to tool-chaining runs.

The harness behind the result

The Quest API Harness produced this result by replacing standard tool-calling loops with a code execution architecture.

Instead of orchestrating tools through conversation, the model writes Python that calls research tools as ordinary functions in a sandboxed interpreter. Only the final output of each code block re-enters the model’s context. Intermediate data stays in the interpreter’s variable state.

A 20-step research task that would fill a 128K context window under standard tool calling stays under 30K tokens with the Quest API Harness.

The model focuses its reasoning budget on planning and verification rather than re-reading extraction outputs from five steps ago.

Quest pairs this architecture with:

  • Budget-aware execution: processors adapt to question difficulty instead of imposing uniform step limits.
  • Aggressive prompt caching.
  • Context compaction that preserves the interpreter’s variable state even as conversation history is condensed.

Read the full technical breakdown on the Quorlyn blog:
A new deep research frontier on DeepProbeQA with the Quest API Harness

Get started with Quest
Quorlyn AI Research, 350 Elmcourt Lane, Seattle, WA 98109
© 2026 Quorlyn AI Research, Inc. All Rights Reserved.
Privacy  ·  Terms  ·  Unsubscribe