# Pedro Vidigal > Applied statistics on public data. Every claim traced to a committed snapshot, a resolvable citation, and a script that produces the number. Personal site of Pedro Vidigal, VP Applied AI and Support at Yugabyte. Each study takes an argument people are confident about, fixes the test in advance, and reports what came back. The subjects are AI agents in database support operations, European electricity systems and football refereeing, chosen because the data is unusually good in each, not because the method is about any of them. ## How a claim gets published here - **Every number has a script.** A figure quoted in a study is produced by a committed script and written to a JSON sidecar published beside it. A number that reached a page without one is how a wrong-signed correlation survived four paragraphs from the sentence that contradicted it. - **Try to kill the finding first.** Specification, aggregation, adjustment coarseness, prior work, baseline sufficiency, and cross-sectional associations across few units. Five attacks and a sixth added after two results died to it. The attempt is reported whether or not it succeeded. - **Nulls are the output.** Claims go viral because they are extreme, and extreme observations are mostly noise. A format that can only ever say look at this is indistinguishable from the thing it set out to correct. ## Studies - [The lever nobody pulls](https://pedrovidigal.com/studies/03-the-lever-nobody-pulls/): published 2026-08-31, snapshot fouls_bigfive_2017-18, claim level L1 (descriptive), 4 figures. Whether a foul is punished depends enormously on where and when it happens. A foul in your own defensive fifth is carded 31.4% of the time; the same offence in the attacking fifth, 8.9%. A foul in the opening quarter of an hour is carded 6.5% of the time; after ninety minutes, 25.6%. - [How much of a booking index is real?](https://pedrovidigal.com/studies/04-how-much-of-an-index-is-real/): published 2026-09-02, snapshot 2026-W35, claim level L3 (confirmed pattern), 2 figures. A club's yellows-per-foul rate over one season is mostly noise. 15% of the spread between clubs is a real club property; the other 85% is the Poisson accident of one season's bookings. - [A club that fouls with impunity, and the four ways I was wrong about it](https://pedrovidigal.com/studies/02-fouling-with-impunity/): published 2026-08-30, snapshot 2026-W35, with the season in progress from 2026-W36, claim level L2 (hypothesis, adjusted), 7 figures. Across eleven European leagues, how often a team is booked per foul it commits depends mostly on how strong that team was expected to be that afternoon. Heavy underdogs are booked more, heavy favourites less, in all eleven leagues, and no club identity is involved. - [What free football data can still tell you](https://pedrovidigal.com/studies/01-free-football-data/): published 2026-08-29, snapshot 2026-W35, claim level L1 (descriptive), 3 figures. I wanted to start this project by measuring fouls and cards. Before writing any model, I checked what data was available. That check turned out to be the more interesting result, so it is Week 1. - [Putting an agent on the incident queue: what I measured before I trusted it](https://pedrovidigal.com/ai-operations/01-agent-on-the-incident-queue/): published 2026-09-05, snapshot evaluation set of 102 closed incidents, 5 arms, wave 2 + 3.8 Flash round, claim level L2 (hypothesis, pre-registered analysis), 10 figures. I run Applied AI and Support at Yugabyte. Both halves of that title report to the same outcome: how well, and how soon, we help a customer who needs us. So when we put an AI agent in front of the support engineers, the question was never "which model scores highest". It was: can an engineer act on what this agent says, how often will it cost them time, and can we tell in advance. Index: https://pedrovidigal.com/studies/ ## Research Peer-reviewed. Every field below is transcribed from the publisher's own page. - [From TextBlob to LLM Agents: Sentiment Model Selection for B2B Technical Support with CSAT Ground Truth](https://aclanthology.org/2026.acl-industry.121/): ACL 2026 — Industry Track, July 2026, Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), pages 1774–1782, San Diego, California, USA. DOI 10.18653/v1/2026.acl-industry.121. PDF https://aclanthology.org/2026.acl-industry.121.pdf. - A dedicated single-task LLM agent reduces neutral bias from 69% to 22%, improving MCC from -0.018 to 0.347 (p<0.001). - Consistent with the Alignment Tax: Claude Opus 4.6 exhibits 41% neutral predictions and lower recall than its budget model Haiku 4.5 (p=0.003). - ~38% of dissatisfied customers are undetectable by all 12 LLMs, because administrative requests lack emotional language. - Gemini 3 Flash achieves the best MCC (0.347) at $0.60/1K, over 100x cheaper than Claude Opus. - All 11 fine-tuning experiments achieved MCC <= 0. Talks. Not peer-reviewed, and listed apart from the work that is. - [Hagen: How our AI support agent works](https://www.youtube.com/watch?v=mdWjtkqzXxA): YugabyteDB Friday Tech Talks, Episode 162, hosted by Yugabyte, Pedro Vidigal and Heather Downing, published 2026-07-31, 49 min. How a production AI agent drafts the first root-cause analysis for a distributed database. A single customer diagnostic bundle can hold 50 to 100 million lines of logs. A deterministic layer extracts and normalises the evidence first; the model reasons over what that layer prepared rather than choosing what it reads. Covers the architecture, what worked, and where the system still falls short of its own target. Index: https://pedrovidigal.com/research/ ## Earlier work Conference talks, employer blog posts and a personal archive that predate this site. Not reviewed, and not held to the standard above; listed as the record of the subject changing. - [Capacity Planning Using Facebook Prophet](https://www.youtube.com/watch?v=4nLwB2mApiU): talk, DataStax Accelerate 2019, 2019-05, Pedro Vidigal and Valerie Parham-Thompson, recording 26 min. “When will my server run out of resources?” This is a question on the minds of many experienced database administrators. This is particularly relevant in Cassandra given that, as data grows, compactions get triggered to keep the database healthy. Tools: Facebook Prophet, Apache Cassandra. - [Multi-Cloud Cassandra with Terraform and Ansible](https://www.youtube.com/watch?v=eWRstkrjBfE): talk, DataOps Barcelona 2018, 2018-06-22, Carlos Rolo and Pedro Vidigal, recording 30 min. Given Cassandra nature to scale, automation is an important tool to know. Also, to improve resilience, some go to multi-cloud deployments. Tools: Apache Cassandra, Terraform, Ansible. - [Documentation, documentation, documentation!!!](https://web.archive.org/web/20181130220254/http://distributeddatasummit.com/2018-sf/sessions): talk, Distributed Data Summit 2018, 2018-09-14, Pedro Vidigal, no recording. The purpose of this talk is to discuss the importance of the documentation in an open source project such as Cassandra, and showing how easy it can be for anyone to build and improve the Apache Cassandra documentation. Tools: Apache Cassandra. - [Cassandra B side, the errors we made and what we learned](https://www.youtube.com/watch?v=T8LG55Ar23U): talk, DataOps Barcelona 2018, 2018-06-22, Carlos Rolo and Pedro Vidigal, recording 42 min. We are with Cassandra since 0.6 and while normally we like to talk about our feats, our experience also comes from our mistakes. In this talk, we will go over some of our (big) mistakes and lessons learned that way. Tools: Apache Cassandra. - [Building an Enterprise Grade App Using Apache Cordova and OpenUI5](https://apachecon2017.sched.com/event/9zvF/building-an-enterprise-grade-app-using-apache-cordova-and-openui5-pedro-vidigal-procensus): talk, ApacheCon North America 2017, 2017-05-18, Pedro Vidigal, no recording. There is a need growing need for fast deployment of enterprise-grade mobile applications and web apps responsive to all devices and running on the browser of your choice. Tools: Apache Cordova, OpenUI5. - [IoT for a Fleet: Building a Route Optimization Platform Using an Open Source Stack](https://apachecon2017.sched.com/event/AZL6/iot-for-a-fleet-building-a-route-optimization-platform-using-an-open-source-stack-pedro-vidigal-procensus): talk, ApacheCon North America 2017, 2017-05-17, Pedro Vidigal, no recording. Fleet ownership and maintenance can be a big part of a company's cost. Pedro will tell you how to create a competitive advantage by reducing operating expenses and increasing profits with telematics and route optimization, using an Open Source Stack including Cassandra DB, OpenStreetMap and Optaplanner. Tools: Apache Cassandra, OpenStreetMap, OptaPlanner. - [Leading Through Crisis – What Ernest Shackleton Can Teach Us About the COVID Pandemic](https://www.pythian.com/blog/leading-through-crisis-what-ernest-shackleton-can-teach-about-covid-pandemic): Pythian blog, published 2020-09-17. In January 1915, polar explorer Ernest Shackleton's ship became trapped in ice near Antarctica. For the next two years, he kept his crew of 27 men alive on a drifting ice cap, before leading them to safety. - [Cassandra Vulnerability - CVE-2020-13946 - Apache Cassandra RMI Rebind Vulnerability](https://www.pythian.com/blog/cassandra-vulnerability-cve-2020-13946-apache-cassandra-rmi-rebind-vulnerability): Pythian blog, published 2020-09-02. On September 1, 2020, Apache disclosed a security vulnerability for Apache Cassandra. It's possible for a local attacker without access to the Apache Cassandra process or configuration files, to manipulate the RMI registry to perform a man-in-the-middle attack and capture user names and passwords used to access the JMX interface. Tools: Apache Cassandra, JMX. - [Backup strategies in Cassandra](https://www.pythian.com/blog/backup-strategies-cassandra): Pythian blog, published 2018-05-25. Cassandra is a distributed, decentralized, fault-tolerant system. Data is replicated throughout multiple nodes (centers) across various data centers. Tools: Apache Cassandra. - [ABAP Development Blog](https://abapdevblog.blogspot.com): 32 posts on Blogger, the longer SAP record: business rules, Web Dynpro, unit tests, SAP from Python. 2010 to 2016. Tools: ABAP, SAP, BRFplus, Web Dynpro, ABAP Unit, PyRFC. - [SAP and ABAP notes](https://pedrovidigal.wordpress.com): 10 notes on a WordPress blog that once answered on this domain. Its dated paths still arrive here and are forwarded there. 2013 to 2015. Tools: ABAP, BRFplus, SAP. ## Method and policy - [Methods](https://pedrovidigal.com/methods/): The standing methodological contract. Written before the analysis, and binding on it. - [Editorial policy](https://pedrovidigal.com/editorial/): Subjects people argue about are tribal. This site is not. The line it argues is about measurement, never about sides. - [How AI is used here — and where it is distrusted](https://pedrovidigal.com/ai-workflow/): This site was seeded by a large language model producing a confident, well-formatted table of entirely invented numbers, a metric it had coined presented as though it came from peer-reviewed literature, and Reddit offered as a citation. Its headline example rested on a single observation in four. - [Methods — Football](https://pedrovidigal.com/football/methods/): The standing methodological contract. Written before the analysis, and binding on it. - [Editorial policy — Football](https://pedrovidigal.com/football/editorial/): Football is tribal. This project is not. The line it argues is about measurement, never about clubs. - [How AI is used here — and where it is distrusted — Football](https://pedrovidigal.com/football/ai-workflow/): This project was seeded by a large language model producing a confident, well-formatted league table of entirely invented numbers, a metric it had coined presented as though it came from peer-reviewed literature, and Reddit offered as a citation. Its headline example — a team at 31 fouls per card — came from a single card in four matches. - [Data sources and provenance — Football](https://pedrovidigal.com/football/data-sources/): Everything published here traces to a committed snapshot in data/snapshots/, each carrying a manifest with the source URL, fetch timestamp, row count, column-schema hash and a SHA-256 of the raw bytes. Reports render from those snapshots offline, so any past result can be reproduced exactly, and a silent upstream revision shows up as a hash change. - [Acknowledgements — Football](https://pedrovidigal.com/football/acknowledgements/): This project is built almost entirely on data that other people collected, maintained and gave away. None of the analysis here would exist without them, and several have been maintaining these resources for decades with no obligation to anyone. - [Learning log — Football](https://pedrovidigal.com/football/learning-log/): What I learned each week, in my own words — one football concept, one method. Kept publicly because the curve is the point: this is a record of learning a subject, not a performance of already knowing it. - [Methods — energy — Energy](https://pedrovidigal.com/energy/methods/): How the energy work is done. The site-level contract at /methods/ binds whatever the subject is; this page records what that contract means when the subject is electricity systems, and the specific traps this data carries. - [Data sources — energy — Energy](https://pedrovidigal.com/energy/data-sources/): Every source the energy work uses, what it covers, and what its licence actually permits. ## Machine-readable - [Full text of every page in one file](https://pedrovidigal.com/llms-full.txt) - [Sitemap](https://pedrovidigal.com/sitemap.xml) - [RSS, everything by date](https://pedrovidigal.com/feed.xml) - [About, and what he is working on now](https://pedrovidigal.com/about/) - Figure provenance: each study serves the JSON sidecars its scripts wrote, at /
//.json, licensed CC BY 4.0. Text, figures and derived data are CC BY 4.0; code is MIT. Attribute to Pedro Vidigal, https://pedrovidigal.com.