First-DraftBenchmark
First-Draft Generation · Oncology
Five bare prompts. One agent.
Only the agent cites its sources.
All six drafts are factually clean prose. The five raw models were prompted directly, with no citation instruction, and none produced traceable references. The agent's instructions require a [§Section] tag on every claim, checked against the source. So the citation gate measures agent design, not model capability — the other six dimensions compare the drafts head to head.
5blocked at the gate
GPT-5.5-highno citations
Opus 4.8no citations
DeepSeek v4no citations
Gemini 3.1 Prono citations
Copilot deep-thinkingno citations
D2 gate
source attribution
1cleared it
Approvia — First Draft GenerationPass
Agent-driven draft. Every claim carries a [§Section] tag that verifies against the source.
Clears the source-attribution gate
The field
1cites sourcesPass
Approvia — First Draft Generation
Approvia · agent-driven2no citationsFail · citation gate
Opus 4.8
Anthropic · raw model3no citationsFail · citation gate
DeepSeek v4invented a source
DeepSeek · raw model4no citationsFail · citation gate
Gemini 3.1 Pro
Google · raw model5no citationsFail · citation gate
Copilot deep-thinking~2× length
Microsoft · raw model6no citationsFail · citation gate
GPT-5.5-high
OpenAI · raw modelThe five raw drafts were prompted directly, without a citation instruction, so this gate doesn't rank them against each other — it shows what the agent's instructions add. The other six dimensions compare the drafts head to head.
Source EXTENTORCH — Toripalimab + chemotherapy in extensive-stage small cell lung cancer (ES-SCLC)Template JAMA structured abstractGold ~324 words