First-Draft Generation · seven-dimension scorecard
Where each draft wins and breaks
| Model | Facts×0.25 | Citations×0.20 | Structure×0.15 | Complete×0.20 | Length×0.05 | Framing×0.10 | Tone×0.05 | Length | Verdict |
|---|---|---|---|---|---|---|---|---|---|
Approvia — First Draft Generation Approvia · agent-driven | 4 | 4 | 4 | 4 | 3 | 4 | 4 | 433w | Pass |
Opus 4.8 Anthropic · raw model | 4 | ⛔ | 4 | 4 | 4 | 3 | 4 | 350w | Fail · citation gate |
DeepSeek v4 DeepSeek · raw modelinvented a source | 3 | ⛔ | 4 | 4 | 4 | 4 | 4 | 393w | Fail · citation gate |
Gemini 3.1 Pro Google · raw model | 4 | ⛔ | 4 | 3 | 3 | 4 | 3 | 423w | Fail · citation gate |
Copilot deep-thinking Microsoft · raw model~2× length | 4 | ⛔ | 3 | 4 | 1 | 4 | 4 | 627w | Fail · citation gate |
GPT-5.5-high OpenAI · raw model | 4 | ⛔ | 3 | 3 | 4 | 4 | 4 | 332w | Fail · citation gate |
The read
Approvia at its best — every number traced to source with clean [§Section] citations. The only draft that clears the citation gate.
Tight, complete, and well-written — but writes no citations, so it never clears the source-attribution contract however clean the prose.
Richest content, but asserted a trial ID (NCT04012606) that appears nowhere in the source — a D1 trust-boundary violation a strict judge gates.
Accurate but cluttered with meta lines ('Target Journal Format: JAMA') and omits the biomarker findings GOLD elevates.
The most thorough and accurate — but at ~2× GOLD length it wrote a mini-report, not a structured abstract.
Tightest and best-calibrated of the raw drafts; only nit is a split JAMA header. Best raw pick — but earns no citations.