Join the conversation
Join the community of Machine Learners and AI enthusiasts.
Sign UpSTE is one of the few style guides you can lint.
Procedural sentences cap at 20 words, descriptive at 25. One instruction per sentence. That makes "did it work" countable, not a vibe.
Have you counted the share of sentences over the cap, before and after? That would separate "follows the rule" from "just sounds plainer".
I transformed a wordy repo:
Size
- Words: 115,559 β 124,979 (+8.2%)
- Sentences: 4,795 β 10,983 (+129%)
Sentence length
- Average words per sentence: 24.1 β 11.4 (β53%)
- 90th-percentile words per sentence: 46 β 20 (β57%)
- Longest sentence: 175 β 62 words (β65%)
- Sentences over 20 words: 53.1% β 8.0% (β85%)
- Sentences over 25 words: 41.6% β 2.5% (β94%
- Sentences over 40 words: 15.8% β 0.2% (β99%)
Paragraphs
- Paragraphs with more than 6 sentences: 1.2%
Readability
- Flesch reading ease: 54.1 β 72.5 (+18 points)
- Flesch-Kincaid grade: 11.7 β 6.0 (β5.7 grad
Vocabulary and grammar
- Unique words: 7,233 β 5,796 (β20%)
- Passive voice per 1k words: 8.5 β 5.7 (β33%
- Complex tenses per 1k words: 1.5 β 0.5 (β66%)
- "-ing" words per 1k words: 26.9 β 12.2 (β55
Punctuation
- Semicolons per 1k words: 9.8 β 0.9 (β91%)
- Em dashes per 1k words: 8.9 β 0.1 (β99%)
Sentences over 25 words 41.6% to 2.5%: that is the STE pass, more or less.
The pair I would watch is sentences +129%, words +8.2%. The long sentences were split, not cut. About 9,400 words went in, the glue a split needs: repeated subjects, "then", "after that". STE's answer to that is the controlled vocabulary, and unique words -20% says the dictionary did some work.
Two things Flesch cannot see. One, whether one instruction per sentence held, which you can count as imperatives per sentence. Two, whether the agent still does the task the same way, since it reads this file every session.
Did task behaviour change on anything you run often? Same prompts, before and after the rewrite, is the only test that matters for a claude.md.
I think STE is to help humans examine model output especially its design, hence you can provide the right feedback
That changes the test. If the reader is the human, the number to count is review time and catch rate, not agent behaviour.
STE was written for maintenance manuals read under time pressure. Same job here: a reviewer skimming a design decision at 2 am.
Did the 11-word version surface anything in the model's reasoning you had missed in the 24-word version?