The 200-account holdout corpus, the exact prompts, the ground-truth labels, and a one-file runner we used to benchmark six frontier LLMs against our churn engine. Run it yourself, then try to beat 0.94 precision at recall ≥ 0.5.
| System | Precision | Recall | AUC |
|---|---|---|---|
| Resonant IQ engine | 0.94 | 0.53 | 0.93 |
| GPT-5 | 0.21 | 0.97 | 0.75 |
| Claude Opus 4.8 | 0.18 | 1.00 | 0.78 |
| Gemini 2.5 Pro | 0.18 | 1.00 | 0.64 |
| DeepSeek | 0.23 | 1.00 | 0.70 |
| Kimi K2 | 0.19 | 1.00 | 0.63 |
| GLM-4.6 | 0.18 | 1.00 | 0.61 |
Every LLM caught almost every churner (recall 0.97–1.00) and was wrong on roughly four out of five accounts it flagged (precision 0.18–0.23). Our engine caught about half the churners it was shown, and was right on 94% of the accounts it raised a hand on. Precision/recall are measured on this case-enriched corpus (32 of 178 scorable accounts were real churners).
It's free and public: download it on GitHub. Prefer it in your inbox, or want the follow-up post with the community's results? Leave your email.