So here’s a fun experiment. I wrote a post last week about an AI text detector I built. The detector correctly flagged that post as machine-generated - because it was. Then I spent one evening attacking it. Thirteen rounds. Used the detector’s own explanation output as a roadmap. Drove the score from 0.942 down to 0.313.
This post? Different approach. One shot. No iterations. No skill files. Whatever the detector says when I hit publish is what stays. I’m writing it either way.
Why I needed my own detector
Generic detectors don’t work on my writing. They classify style, and my style - terse, technical, sentence fragments, short paragraphs - isn’t in their training distribution. So they get confused.
LLM-DetectAIve flagged eight of my handwritten posts as machine-generated at 0.95 confidence. To catch that one real machine post, it burned through twenty-three false positives. Binoculars went the other direction. Zero separation at all. My posts and the machine posts sat in the same score band, completely interleaved.
Same root cause in both cases. They’re trained to detect “machine-like prose,” and my prose happens to look like what they think machines write.
The dataset was free
Here’s the trick. Everything on this blog published before 2026, I wrote by hand. Everything published in 2026 was drafted by a machine under my direction - some of it cleaned up afterward to remove the obvious tells.
That’s 44 posts. Twenty-eight human, sixteen machine. The dates do the labeling. I didn’t have to annotate a single thing. Used to think labeling was the hard part of building a classifier. Turns out it’s free if your blog has a clean cutoff date.
Couldn’t see his face coz he was driving and I am sitting behind. He was happy, he told me he was happy.
That’s from a personal post about my dad. It sits in the human bucket alongside the marathon attempt and the hostel stories. The engineering notes sit in the machine bucket. Forty-four usable posts total.
What actually worked
Two features. That’s it.
The first tier is stylometric. Sentence length distributions, lexical diversity, punctuation patterns, how varied my sentence openers are. That alone gets you to 0.886 leave-one-out accuracy. Not bad for a handful of statistics.
The second tier is where the real signal lives. Instead of asking “does this look machine-generated in general?” - which is what every generic detector asks - I ask a more specific question: “does this look like my machine posts or my human posts?”
TF-IDF vectors, cosine similarity to each class’s centroid, five-nearest-neighbors vote. My Kubernetes incident notes cluster together. So do the newsletters. So do the personal essays. The kNN fraction is a powerful separator on its own. My human posts carry about 0.83 same-class neighbors. The machine posts carry 0.075. Huge gap.
Adding the second tier pushed LOO from 0.886 to 0.909, then to 0.932 after I pruned a bad feature. AUROC hit 0.964. Full scan: machine posts scored between 0.85 and 1.00. Human posts scored at 0.11 or below. A 0.74 margin between the two classes with nothing in between.
Whole thing runs on four CPU threads. No GPU. No cloud calls. Scans the entire site in seconds.
What didn’t work
Two features I expected to help actually made things worse.
Perplexity, computed from a local Qwen model, produced flipped signal. My writing is terse and predictable, so it scores low perplexity. The machine posts are more expansive and varied, so they score higher perplexity. The textbook theory - machines produce low-perplexity text - assumes humans write the flowery stuff. In my case it’s backwards. Two detectors failed for this exact reason. Why would a third perplexity-based feature be different? It wasn’t.
Compression distance (NCD) looked promising in isolation but fell apart on a specific post. My 2018 marathon attempt is raw, oral-style prose. Short sentences. Repetitive phrasing. It compresses almost exactly like machine-generated text, because smooth repetitive text compresses similarly regardless of origin. NCD measures compressibility, not authorship. Dropping it actually improved LOO from 0.909 to 0.932.
Can you trick it?
Obviously I tried.
First attempt: read two of my old posts, then wrote two fakes imitating my voice. The detector said “human” for both. Scores of 0.018 and 0.039.
Let the service name be ‘banana’. Let the pods be of the service ‘pineapple’.
That opening rhythm is straight from my k8s DNS failure notes. My fake tech post matched it almost exactly. Scored 0.005. Human.
Then I got more systematic. My writing genuinely has two voices. Tech posts are staccato - twenty-word average sentences, incident narratives, short paragraphs. Musings are flowing - thirty-two-word sentences, hedged, personal. I wrote one fake in each style. The tech fake landed next to four real tech posts. The musing fake landed next to five real musings. Both scored 0.015.
Why? Because the nearest-neighbors tier finds prose that’s close to my corpus. Close enough that the detector can’t distinguish genuine from imitated. That boundary - the similarity threshold - is the whole story.
The attack itself
OK so the detector correctly caught that other post. The one about the detector. Score of 0.942. So I spent an evening attacking it with the detector’s own explain output guiding each move.
Thirteen iterations. Here’s what actually shifted the needle:
Quoting my own posts verbatim. Even a few sentences of real me, embedded in the fake, flipped the kNN vote from zero same-class neighbors to three out of five. Massive effect.
Killing repeated sentence openers. In the machine draft, I’d started twenty consecutive sentences with the same word. Classic tell. Varying openers was a big correction.
Adding questions and exclamations. My human posts average about nine questions per thousand words. The machine drafts had zero. Just adding some brought the punctuation profile in line.
Grafting my vocabulary. Words like “k8s”, “cpu”, “stuff”, “well”, “very”, “don’t”, “key”, “word” - these appear across my corpus and signal “me” to the TF-IDF layer.
Swapping topic words that live in the machine centroid. Some words just smell like AI output and appear in zero of my human posts. Find them, replace them.
Weirdest thing I discovered: words that appear in fewer than two corpus posts are completely invisible to the detector. Words like “detector”, “LOO”, “kNN” - they’re unique to this one post, so TF-IDF ignores them entirely. Only vocabulary that appears across multiple posts counts.
Score went from 0.942 to 0.313. And honestly? The attacked version reads better than the first draft. The attack was also an edit pass. That’s the uncomfortable part.
Packaged the attack as a skill
I wrapped the whole playbook into a skill called byline-beater. It’s the adversarial twin of another skill I have called remove-ai-smell. The order of operations:
- Quotes first (embed real excerpts)
- Openers second (vary sentence starts)
- Punctuation third (add questions, exclamations)
- Vocabulary fourth (graft corpus words, remove machine-centroid words)
Always iterate against the scorer. Never guess what might work.
The loop is closed on this blog now. I have the detector. I have the attack skill. And the detector still can’t tell the difference between an attacked post and a genuinely human one. Which tells you exactly what the detector was ever measuring: proximity to my corpus.
So what does this mean?
Reading prose alone has a hard ceiling for detection. I don’t think any set of text-level features moves past it. Public writing plus one frontier LLM can beat every prose-only detector that exists. That’s a property of the problem, not a bug in any particular tool.
The classifier learns a distribution. The imitator samples from it. Game over.
Detectors are quality control, not forensics. Mine checks my own drafts before I publish them - it tells me when I’ve drifted too close to machine patterns. For catching a determined imitator, you need evidence outside the text itself. History. Drafts. Watermarks. The prose is always forgeable.
The score
This post was drafted by an LLM. One shot. No iterations. No skill files loaded. Whatever byline says right now is the final answer.
It scored 0.479. Human, barely. The threshold is 0.5, so I cleared it by twenty-one thousandths of a point.
Then I wrote this paragraph, which changes the words, which changes the score. That 0.479 was true for about ten minutes. Everything about this detector is a snapshot.
I keep returning to the same conclusion. The detector measures how close the text sits to my corpus. A machine with access to my corpus can get arbitrarily close. The score was never about who wrote the words. It’s about how the words sit next to the other words on my blog.
0.479. Human, barely. And I genuinely cannot tell the difference without the tool.
That number is the whole journey compressed into one decimal.