The Quest API Harness
The Quest API Harness is what produced the result.
The Quest API Harness replaces standard tool-calling loops with a code execution architecture. Instead of the model orchestrating tools through conversation, it writes Python that calls research tools as ordinary functions in a sandboxed interpreter. Only the final output of each code block re-enters the model’s context. Intermediate data stays in the interpreter’s variable state.
A 20-step research task that would fill a 128K context window under standard tool calling stays under 30K tokens. The model focuses its reasoning budget on planning and verification rather than re-reading extraction outputs from five steps ago.
Quest pairs this with budget-aware execution, where processors adapt to question difficulty rather than imposing uniform step limits, aggressive prompt caching, and context compaction that preserves the interpreter’s variable state even as conversation history is condensed.
|