
Does a Polite AI Actually Win More Users?
TL;DR & Core Takeaways
Tone is a tie-breaker, not a cure-all: A warmer, more positive tone increases an AI's win rate—but almost exclusively when two models are evenly matched on functional accuracy.
Accuracy beats friendliness every time: When a model gives incorrect or hallucinated answers, no amount of enthusiastic or polite phrasing will save its user preference score.
Keep prompts consistent: Users expect equal tone across coding and general chat tasks; there’s no need to build overly strict "neutral tone" rules for technical domains.
The Question
LLM teams spend significant compute and prompt-tuning effort refining model personality and tone. But does being "nicer" actually make users prefer one model over another?
To find out, I analyzed 50,000+ head-to-head LLM battles from the LMSYS Chatbot Arena dataset. Using DuckDB and logistic regression, I evaluated whether positive sentiment acts as a primary driver of user preference or simply a secondary cosmetic layer.
Key Insight 1: The Parity Curve
When functional capabilities are equal, tone takes over. When two models provide equally accurate answers, the positive sentiment delta produces its highest marginal win rate gain (peaking right at $P \approx 0.50$). This highlighted "Tie-Breaker Zone" shows where prompt polish delivers real ROI—helping an otherwise identical response edge out the competition.

Key Insight 2: Capability Asymmetry
Politeness cannot bridge an accuracy gap. Comparing a top-tier model against a struggling one reveals the limits of tone. The red curve demonstrates that even if an inferior model adopts a wildly positive, helpful tone, its overall win probability stays anchored near the bottom. Product teams should always prioritize hallucination reduction and core accuracy over tone adjustment.

Methodology Highlights
Data Processing: Ingested and structured 50,000+ paired preference prompts using DuckDB.
Statistical Modeling: Built binary logistic regression models in statsmodels to isolate sentiment coefficients while controlling for task domain and response length.
Data Visualization: Generated custom SVG visuals in Seaborn and Matplotlib, styled for executive presentation.
Want to See the Code?
Interested in the full data pipeline, statistical tests, or raw Jupyter notebook?