Churn Benchmark

Get the churn benchmark pack

The synthetic 200-account corpus, the exact prompts, the ground-truth labels, and a one-file runner we used to benchmark six frontier LLMs against our churn engine. Run it yourself, then try to beat 0.94 precision at recall ≥ 0.5.

The results

The corpus is synthetic — 200 B2B SaaS accounts built by our own simulator, not real customer data — and we publish it with the prompts and labels so you can check the numbers yourself.

System Precision Recall AUC
Resonant IQ engine 0.94 0.53 0.93
GPT-5 0.21 0.97 0.75
Claude Opus 4.8 0.18 1.00 0.78
Gemini 2.5 Pro 0.18 1.00 0.64
DeepSeek 0.23 1.00 0.70
Kimi K2 0.19 1.00 0.63
GLM-4.6 0.18 1.00 0.61

Every LLM caught almost every churner (recall 0.97–1.00) and was wrong on roughly four out of five accounts it flagged (precision 0.18–0.23). Our engine caught about half the churners it was shown, and was right on 94% of the accounts it raised a hand on. Precision/recall are measured on this case-enriched corpus (32 of 178 scorable accounts were labeled churners).

Get the pack

It's free and public: download it on GitHub. Prefer it in your inbox, or want the follow-up post with the community's results? Leave your email.

We'll email you the download link now, plus the follow-up post and occasional updates when we publish new benchmark results or engine deep-dives. No cadence promises, easy unsubscribe.