First-DraftBenchmark

First-Draft Generation · Oncology

Five bare prompts. One agent.
Only the agent cites its sources.

All six drafts are factually clean prose. The five raw models were prompted directly, with no citation instruction, and none produced traceable references. The agent's instructions require a [§Section] tag on every claim, checked against the source. So the citation gate measures agent design, not model capability — the other six dimensions compare the drafts head to head.

5blocked at the gate
GPT-5.5-highno citations
Opus 4.8no citations
DeepSeek v4no citations
Gemini 3.1 Prono citations
Copilot deep-thinkingno citations
D2 gate
source attribution
1cleared it
Approvia — First Draft GenerationPass

Agent-driven draft. Every claim carries a [§Section] tag that verifies against the source.

Clears the source-attribution gate

The field

1
Approvia — First Draft Generation
Approvia · agent-driven
2
Opus 4.8
Anthropic · raw model
3
DeepSeek v4invented a source
DeepSeek · raw model
4
Gemini 3.1 Pro
Google · raw model
5
Copilot deep-thinking~2× length
Microsoft · raw model
6
GPT-5.5-high
OpenAI · raw model

The five raw drafts were prompted directly, without a citation instruction, so this gate doesn't rank them against each other — it shows what the agent's instructions add. The other six dimensions compare the drafts head to head.

Source EXTENTORCHToripalimab + chemotherapy in extensive-stage small cell lung cancer (ES-SCLC)Template JAMA structured abstractGold ~324 words