An LLM Judge That Watches for the Sighs: Phoenix-Evals Ships a 'User Friction' Evaluator

A new evaluator scores the implicit signals users send when a product isn't working — the frustration you can feel but can't easily measure.

Most evals grade whether a model produced the right answer. The new "user_friction" evaluator shipped in phoenix-evals 3.2.0 grades something harder: whether the user was quietly struggling. As @ehutt_ described it, the evaluator is an LLM judge that watches for the implicit feedback signals users send when something isn't working — the rephrased questions, the abandoned threads, the mounting exasperation that never becomes an explicit complaint.

Unlock the full briefing

Get every story in today's briefing, the full archive, and the daily AI intelligence brief.

All stories today

Full archive

Daily brief

Cancel anytime. Payments powered by Stripe.