prosemeter

Research

prosemeter was built to answer one question — does a writing instruction actually change how a model writes? — and then used to answer it. 596 generated answers across 7 runs, every score committed, every answer kept.

These are the conclusions. Each links to the study that established it. Where a study found the opposite of what it set out to show, it says so; three of the six did.

What instruction wording moves

Concrete rewrite rules cut length by about a third

Measured within a single run against a no-instruction control on the same tasks: 438 words down to 298, a 32% cut. Jargon fell 2.1 points and sentence-simplicity rose 19.7 points.

What the instruction is actually worth, against no instruction at all →

Naming a technique beats spelling it out

Replacing one rule with its label — Use BLUF: bottom line up front — and holding the other four fixed cut 38% off the control against the spelled-out sentence's 31%, while tying it on jargon, sentence-simplicity, grade band and clarity.

Naming a technique instructs better than spelling it out →

Naming the better word works; asking for fewer does not

“Avoid jargon” does not reduce jargon. Naming the swap — use over utilize, start over initiate — does. The vague instruction lands most of the length win and almost none of the vocabulary win.

Seven instructions, four tasks, and the one that won →

What it does not move

Factual accuracy, at all

Two runs and one targeted mechanism test found the same thing: wording moves length and vocabulary two to three times over, and moves correctness not at all. The two highest-scoring instructions tied the control for worst accuracy. Task difficulty dominated — accuracy spanned 3/15 to 14/15 by task against 15/30 to 21/30 by instruction.

A high score is not a correctness signal. Verify facts separately.

Does cutting words cut the qualifiers that carry meaning? →

The writing, when you revise toward the score

Showing a reviser prosemeter's own findings beat a blind “just revise it” pass on 29 of 30 drafts, worth 4.3 composite points. Then a blind reader picked the blind revision 18 times to 9, and 14 to 4 on the explain-a-concept tasks.

The gain never left the dimensions the findings came from. Across the five dimensions no finding pointed at, the two arms differ by at most 0.4 points in either direction. The loop moves what it measures.

The revise loop raises the score without improving the writing →

Whether an explanation is legible

A reply scoring 85 drew “I have no idea what you're saying”; the rewrite that landed scored 84. An arm built to fix that failure moved the composite by half a point. The instrument cannot see the thing, so it cannot score the fix — it can only price it.

Naming a technique instructs better than spelling it out →

What the method taught us

Read the dimensions, not the composite

The composite is a weighted average, so it averages the real effects away. Run 1 spanned 79.5–86.6 across seven instructions while the spread inside a single instruction was about 15 points — so no composite gap in it meant anything. The dimensions moved two to three times over the same answers.

Seven instructions, four tasks, and the one that won →

Model judges read position before they read prose

Three independent judges over the same 30 pairs agreed with each other 64% of the time, a Fleiss' kappa of 0.30. The cause was presentation order: a judge shown the winning answer first picked it 97% of the time, and 57% when it came second.

Judge every pair both ways and keep only the verdicts that survive, and the picture sharpens rather than fading — 45 of 47 surviving verdicts go the same way. Half the disagreement was the layout.

The judges were reading position, not prose →

A metric that works can be an artifact of your task set

Word count was the sharpest instruction dial on six tasks that were all the same shape, and stopped working on ten mixed ones — natural length varies far more when the registers differ. Sentence-simplicity held across both and is the one to lean on.

The best metric was an artifact of the task set →

Findings are model-specific, and the ranking travels further than the size

Re-run on Sonnet 5, the winning instruction still won, but moved length 12% where it moved Opus 41% — its uninstructed output was already a third shorter, so there was less to win.

Does the winning instruction survive a change of model? →

A dimension that never fires is not thereby useless

Grade band scores a perfect 100 on 87 of every 100 answers in this corpus, which looked like dead weight. It is not: the corpus is answers written by a model asked to be clear, so it contains almost nothing for a ceiling to catch. Removing the ceiling let a deliberately-jargon-heavy fixture score 100 at reading grade 25.8.

The studies

The data

Every answer and every score is in the repository. 596 answers across 7 runs, on claude-opus-5 and claude-sonnet-5, with 20 carrying a reviewer's finding of a specific technical error.

Those errors are the point, not a defect. The finding that wording does not reach accuracy can only be shown by keeping the wrong answers beside the right ones, with the instruction that produced each. Every file says in its own front matter that it is experiment output and was never fact-checked as documentation.