Skip to content

How copywriting evaluation works

The problem: “is this article good?” hides two independent questions — is it true and is it worth reading — and optimizing either alone produces failure: accurate-but-unreadable, or delightful fabrication. The platform’s doctrine makes them two co-equal gates: an article ships only in the top-right quadrant.

ShipsRewrite styleRejectFix the factsHard to readEasy and pleasantFabricatedTrue and accurateThe two co-equal gates

Fabrication hunting is adversarial: strip every claim, then try to ground each one. Four grounding classes are accepted — a real source URL from the article’s own research bundle, persona facts from the author’s bio, brand specifics from the company’s declared profile, and reviewer-confirmable textbook knowledge of the niche. Everything else is a fabrication finding. The class list exists because each one was once a false positive that burned review time — codifying them made the hunt repeatable.

Readability is judged by a structured editor-grade reading rubric — an actual read-through scored on defined axes, not a proxy metric. It is deliberately not an LLM score: cheap LLM judges measurably anti-correlated with human quality ratings on this workload (see Quality gates), so the reading gate stayed human- calibrated. An anti-template battery backs it mechanically: banned AI-vocabulary replacements, banned heading shapes (“Core concepts behind X…”), and a corpus-level twin-similarity pass that catches articles converging on the same skeleton.

Layer 3 — evidence-anchored judging, where LLM opinion is used at all

Section titled “Layer 3 — evidence-anchored judging, where LLM opinion is used at all”

Where an LLM verdict participates (as advisory signal, never as the gate), it is structurally constrained: a locked 0/1/2 rubric per category, a verbatim quote required for any non-zero score — checked mechanically as a substring — multiple samples with a median, and a judge from a different model family than the writer. A judge that cannot cite cannot praise. The full formal judge (with human-set calibration) remains an analytics instrument: its calibration study was never completed, so it never earned gate status — a deliberate application of the platform’s own “deterministic decides, LLM advises” doctrine to itself.

Every published article carries a quality stamp computed from the deterministic automatic categories (structure, AI-marker absence, technical SEO) minus defect penalties — brand-safety hits flip the verdict, not the number, so a beautiful article that promotes the wrong brand fails loudly rather than averaging out. The designed next phase — a “rising tide” that re-enriches old articles as the research corpus matures — is on the roadmap, not yet built.

Human-calibrated reading gates do not scale like a metric — that is the accepted cost of the measured anti-correlation result. And the two-gate doctrine means an article can loop: fixed facts can degrade flow, style rewrites can drift facts; the loop converges because each gate re-checks after the other’s edit, but convergence costs iterations.

Any article’s quality stamp and signal battery are readable through the API (quality domain); the anti-template ban lists live in the generation pipeline’s version directory — versioned like everything else.

Specs: SPEC-074 (two-gate loop + playbook), SPEC-065 (evidence-anchored judge), SPEC-106 (post-publish scoring; rising tide designed).