Cover image for Does a Polite AI Actually Win More Users?

Does a Polite AI Actually Win More Users?

TL;DR & Core Takeaways

Tone is a tie-breaker, not a cure-all: A warmer, more positive tone increases an AI's win rate—but almost exclusively when two models are evenly matched on functional accuracy.

Accuracy beats friendliness every time: When a model gives incorrect or hallucinated answers, no amount of enthusiastic or polite phrasing will save its user preference score.

Keep prompts consistent: Users expect equal tone across coding and general chat tasks; there’s no need to build overly strict "neutral tone" rules for technical domains.


The Question

LLM teams spend significant compute and prompt-tuning effort refining model personality and tone. But does being "nicer" actually make users prefer one model over another?

To find out, I analyzed 50,000+ head-to-head LLM battles from the LMSYS Chatbot Arena dataset. Using DuckDB and logistic regression, I evaluated whether positive sentiment acts as a primary driver of user preference or simply a secondary cosmetic layer.

Key Insight 1: The Parity Curve

When functional capabilities are equal, tone takes over. When two models provide equally accurate answers, the positive sentiment delta produces its highest marginal win rate gain (peaking right at $P \approx 0.50$). This highlighted "Tie-Breaker Zone" shows where prompt polish delivers real ROI—helping an otherwise identical response edge out the competition.

Plot showing the parity curve

Key Insight 2: Capability Asymmetry

Politeness cannot bridge an accuracy gap. Comparing a top-tier model against a struggling one reveals the limits of tone. The red curve demonstrates that even if an inferior model adopts a wildly positive, helpful tone, its overall win probability stays anchored near the bottom. Product teams should always prioritize hallucination reduction and core accuracy over tone adjustment.

Plot showing the parity curve


Methodology Highlights

Data Processing: Ingested and structured 50,000+ paired preference prompts using DuckDB.

Statistical Modeling: Built binary logistic regression models in statsmodels to isolate sentiment coefficients while controlling for task domain and response length.

Data Visualization: Generated custom SVG visuals in Seaborn and Matplotlib, styled for executive presentation.

Want to See the Code?

Interested in the full data pipeline, statistical tests, or raw Jupyter notebook?

👉 View the Complete Project Repository on GitHub