Skip to content

1 peer-reviewed · 1 talk

Research

Work reviewed or hosted elsewhere. Every figure below is transcribed from the publisher’s own page rather than restated, so a number here can be checked against the source in one click.

Peer-reviewed

  • ACL 2026 — Industry Track ·

    From TextBlob to LLM Agents: Sentiment Model Selection for B2B Technical Support with CSAT Ground Truth

    Pedro Vidigal · pages 1774–1782 · San Diego, California, USA

    A five-year case study of sentiment model selection for customer satisfaction (CSAT) prediction in B2B technical support. The evaluation uses the complete population of CSAT-rated tickets from an enterprise software company: over 500 tickets comprising ~2,500 customer comments from 100+ organizations over five years. 17 approaches across 5 paradigms, plus 11 fine-tuning experiments.

    Findings

    • A dedicated single-task LLM agent reduces neutral bias from 69% to 22%, improving MCC from -0.018 to 0.347 (p<0.001).
    • Consistent with the Alignment Tax: Claude Opus 4.6 exhibits 41% neutral predictions and lower recall than its budget model Haiku 4.5 (p=0.003).
    • ~38% of dissatisfied customers are undetectable by all 12 LLMs, because administrative requests lack emotional language.
    • Gemini 3 Flash achieves the best MCC (0.347) at $0.60/1K, over 100x cheaper than Claude Opus.
    • All 11 fine-tuning experiments achieved MCC <= 0.

Talks

Not peer-reviewed, and listed apart from the work that is.

  • YugabyteDB Friday Tech Talks, Episode 162 · · 49 min

    Hagen: How our AI support agent works

    Pedro Vidigal, Heather Downing · Yugabyte

    How a production AI agent drafts the first root-cause analysis for a distributed database. A single customer diagnostic bundle can hold 50 to 100 million lines of logs. A deterministic layer extracts and normalises the evidence first; the model reasons over what that layer prepared rather than choosing what it reads. Covers the architecture, what worked, and where the system still falls short of its own target.

    Watch on YouTube, 49 min