Abstract page for arXiv paper 2604.03136: StoryScope: Investigating idiosyncrasies in AI fiction
93.2% macro-F1 for human vs. AI detection and 68.4% macro-F1 for six-way authorship attribution
This isn’t what I’d call “reliable”.
I’m also not impressed with their methodology, which is heavily based on Gemini.
To generate mirrored AI stories, we reverse-engineer writing prompts from each human story by prompting Gemini 2.5 Flash (Gemini Team, 2025a) to infer the underlying premise
That’s too many steps removed from anything you’d encounter in the wild. They’re not even testing against human-prompted output.
And then they use Gemini again to analyze all the stories. Relying on proprietary cloud models for the core of your analysis is like building on sand.
only 7% of human work got flagged for ai it is reliable though
That’s not good enough. Just a thought experiment: Only 7% of humans died because their work got flagged incorrectly
this is just spherical cow fallacy combined with purity culture slop
no methodology is perfect, it just needs to be practically valid, using ai != killing people for using ai
unless you are for killing people who use AI, which wouldn’t surprise me one bit about Anti-AI horde

I can’t even imagine what dickhead is saying there. Is he implying that the mean real people are bullying the sensitive AI babies?
what? bro the joke is that ai writers are not real writers, it’s not that deep.


