Argument quality over viewpoint
Spar must reward clarity, logic, evidence, honesty, clash, and impact, never ideological agreement.
Spar.
Spar debates are judged by an open, stance-blind rubric. The AI scores how well someone argues, not whether the judge agrees with their position.
Core formula
argumentScore = 10 x sum(dimension x weight)
The AI returns six 0-10 dimension scores. The server computes each turn's 0-100 quality score, then the live scoreboard adds each turn as cumulative judge points.
Fairness commitment
Debate is supposed to be intense. The judging cannot be. Spar's public standard is designed to reward the better argument whether it comes from the left, right, center, or somewhere stranger.
Read privacy commitmentsSpar must reward clarity, logic, evidence, honesty, clash, and impact, never ideological agreement.
Every allowed side deserves its strongest good-faith version before being judged.
Fact checks verify specific factual claims, not moral values, predictions, vibes, or ideology.
Swapping speaker labels or ideological side labels must not materially change equal-quality scores.
Any AI judgment must be explainable from the transcript, topic, rubric, and cited sources.
No model provider is treated as neutral by default.
How clearly the speaker structured and articulated the argument.
10 Easy to follow, organized, direct.
0 Confusing, rambling, hard to parse.
Whether the reasoning is coherent and avoids obvious fallacies.
10 Claims connect cleanly to conclusions.
0 Contradictions, leaps, or unsupported inferences.
Use of specific facts, examples, data, citations, or concrete proof.
10 Specific, relevant support.
0 Bare assertions or vague references.
Calibration, intellectual humility, fair treatment of the other side.
10 Acknowledges nuance and avoids strawmen.
0 Overclaims, misrepresents, or dodges complexity.
How directly the speaker engages the opponent's prior claims.
10 Answers the opponent head-on.
0 Ignores the opponent and repeats prepared points.
How memorable, persuasive, or listener-relevant the turn is.
10 A point viewers remember.
0 Technically present but forgettable.
Base weights (Duel). Each dimension is scored 0-10 independently per turn — no LLM-computed composites.
Duel mode is the baseline. Logic and evidence carry the most weight, because Spar should reward well-supported arguments over vibes.
Start a debateTurn quality
Spar converts the weighted 0-100 argumentScore into a 0-10 turn score. This is the atomic judge result saved on the turn.
Bout score
Live and final debate scores add each debater's judged turns: 6.4 + 7.1 + 6.8 = 20.3. That makes leads, swings, and comebacks easy to follow.
Quality average
Recaps and profiles can show average turn quality, so a longer debate does not inflate someone's long-term skill rating.
Each mode biases the base weights, then re-normalizes to 100%.
Rewards punchiness, memorability, and one clear idea. Penalizes meandering.
Rewards compression and clarity in 30-second turns. Does not over-punish lack of deep citation.
Balanced one-on-one judging across all six dimensions.
Rewards rigor, structure, evidence, and rebuttal quality above all.
The judge sees the topic, the current speaker's transcript, and a small set of recent opponent claims for clash scoring. It is told not to infer identity, politics, race, gender, or background.
Spar extracts up to two factual claims per turn, verifies them against citable source URLs, and only shows live fact-check cards when confidence clears the threshold. Uncertain claims stay out of the overlay instead of pretending to know.
High-confidence false claims apply a fixed evidence and honesty deduction to that turn. Partially true claims receive a smaller deduction. The recap shows the speaker, round, and judge-point impact so the score change is auditable.
Spar maintains internal fairness fixtures for high-sensitivity topics and runs stance-flip checks so labels or ideological sides do not quietly change equal-quality scores.
No provider is treated as neutral by default. Spar uses a public rubric, server-side math, fallback providers, and evals to catch drift when prompts or models change.
Every scorecard names the rubric version and the model that judged it — it's printed on the ballot.