How copywriting evaluation works
The problem: “is this article good?” hides two independent questions — is it true and is it worth reading — and optimizing either alone produces failure: accurate-but-unreadable, or delightful fabrication. The platform’s doctrine makes them two co-equal gates: an article ships only in the top-right quadrant.
Layer 1 — the truth gate
Section titled “Layer 1 — the truth gate”Fabrication hunting is adversarial: strip every claim, then try to ground each one. Four grounding classes are accepted — a real source URL from the article’s own research bundle, persona facts from the author’s bio, brand specifics from the company’s declared profile, and reviewer-confirmable textbook knowledge of the niche. Everything else is a fabrication finding. The class list exists because each one was once a false positive that burned review time — codifying them made the hunt repeatable.
Layer 2 — the readability gate
Section titled “Layer 2 — the readability gate”Readability is judged by a structured editor-grade reading rubric — an actual read-through scored on defined axes, not a proxy metric. It is deliberately not an LLM score: cheap LLM judges measurably anti-correlated with human quality ratings on this workload (see Quality gates), so the reading gate stayed human- calibrated. An anti-template battery backs it mechanically: banned AI-vocabulary replacements, banned heading shapes (“Core concepts behind X…”), and a corpus-level twin-similarity pass that catches articles converging on the same skeleton.
Layer 3 — evidence-anchored judging, where LLM opinion is used at all
Section titled “Layer 3 — evidence-anchored judging, where LLM opinion is used at all”Where an LLM verdict participates (as advisory signal, never as the gate), it is structurally constrained: a locked 0/1/2 rubric per category, a verbatim quote required for any non-zero score — checked mechanically as a substring — multiple samples with a median, and a judge from a different model family than the writer. A judge that cannot cite cannot praise. The full formal judge (with human-set calibration) remains an analytics instrument: its calibration study was never completed, so it never earned gate status — a deliberate application of the platform’s own “deterministic decides, LLM advises” doctrine to itself.
Layer 4 — after publication
Section titled “Layer 4 — after publication”Every published article carries a quality stamp computed from the deterministic automatic categories (structure, AI-marker absence, technical SEO) minus defect penalties — brand-safety hits flip the verdict, not the number, so a beautiful article that promotes the wrong brand fails loudly rather than averaging out. The designed next phase — a “rising tide” that re-enriches old articles as the research corpus matures — is on the roadmap, not yet built.
The trade-offs, honestly
Section titled “The trade-offs, honestly”Human-calibrated reading gates do not scale like a metric — that is the accepted cost of the measured anti-correlation result. And the two-gate doctrine means an article can loop: fixed facts can degrade flow, style rewrites can drift facts; the loop converges because each gate re-checks after the other’s edit, but convergence costs iterations.
See it in two minutes
Section titled “See it in two minutes”Any article’s quality stamp and signal battery are readable through the
API (quality domain); the anti-template ban lists live in the
generation pipeline’s version directory — versioned like everything else.
Specs: SPEC-074 (two-gate loop + playbook), SPEC-065 (evidence-anchored judge), SPEC-106 (post-publish scoring; rising tide designed).