I agree with one thing you caught: the AI may have...
Tạo vào: 16 tháng 7, 2026
Trả lời bằng GPT-5.6 Thinking bởi Chat01
Tạo vào: 16 tháng 7, 2026
Trả lời bằng GPT-5.6 Thinking bởi Chat01
I agree with one thing you caught: the AI may have “fixed the example” instead of fixing the engine. That’s a common failure mode.
The danger signs are:
If that’s what happened, it simply changed the displayed projection instead of repairing the probability model.
I’d tell the AI something like this:
⸻
TENNISLOCKS ENGINE-WIDE TOTAL GAMES, PLAYER GAMES, AND SCORE DISTRIBUTION REBUILD
This is not a request to fix one match, one tournament, one player, or one example output.
The recent changes corrected several presentation issues, but the remaining problems suggest the canonical probability distribution itself may still be structurally biased.
Do not tune the engine until the displayed example “looks right.”
Do not hardcode projections that happen to match this one match.
Do not introduce special handling for favorites, underdogs, two-set matches, three-set matches, specific scorelines, tours, surfaces, or individual players.
The objective is to improve the mathematical model so that every future match naturally produces coherent probabilities without any match-specific logic.
⸻
First, perform a full audit to ensure no recent change accidentally replaced a probabilistic calculation with deterministic display logic.
Search for code that:
Every displayed statistic must remain a consequence of the underlying probability distribution, not a manually corrected value.
⸻
Treat the exact-score PMF as the foundation of every downstream market.
For every legal BO3 and BO5 scoreline, compute and expose internally:
Verify that every market is generated exclusively from these outcomes.
No downstream market should estimate values independently once the score PMF exists.
⸻
Do not stop at reporting:
Inspect the complete distribution.
Measure:
Determine whether the distribution consistently allocates excessive probability to long matches.
If the right tail is inflated, identify exactly which scorelines contribute the excess probability.
Repair the transition probabilities creating that inflation rather than suppressing the tail afterward.
⸻
Generate frequency tables for every legal scoreline.
Examples:
Determine whether the engine systematically:
Do not compensate by lowering totals afterward.
Repair the score generator itself.
⸻
Investigate every transition that extends a set beyond six games.
Specifically examine transitions from:
Determine whether:
Even small transition errors compound into substantial total-games inflation.
⸻
Do not evaluate only the overall percentage of three-set matches.
Instead determine why the engine enters deciding sets.
For every simulated third set, record:
Identify whether one subsystem systematically keeps matches alive longer than the underlying player strengths justify.
⸻
Winning probability and winning margin are different quantities.
A player projected to win 62% of matches should not automatically produce highly competitive scorelines.
Determine whether the engine compresses victory margins.
Measure:
If favorites consistently win by margins smaller than justified by the underlying hold and break probabilities, total games and loser games will both become inflated.
Correct the margin generation rather than modifying totals afterward.
⸻
Player Games must be computed from the complete exact-score PMF.
Do not derive Games Won from:
For every sportsbook line, compute the probability directly from the PMF:
P(Player Games > Line) equals the sum of probabilities of all scorelines where the player exceeds that line.
Verify that the displayed mean, median, mode, and Over/Under probabilities all originate from the same PMF.
⸻
Compute Over/Under probabilities only by summing the exact-score PMF.
Do not allow:
to modify the final Over/Under probability after the PMF is finalized.
Those systems may diagnose or validate the PMF but must not override it.
⸻
For every simulation verify that:
Winner ⇄ Exact Score
Winner ⇄ Sets Won
Winner ⇄ Player Games
Winner ⇄ Total Games
Match Length ⇄ Exact Score
Match Length ⇄ Total Games
Match Length ⇄ Tiebreaks
Exact Score ⇄ Player Games
Exact Score ⇄ Sets Won
Exact Score ⇄ Total Games
Exact Score ⇄ First Set
Total Games ⇄ Player Games
Props ⇄ Canonical PMF
Any disagreement indicates an architectural bug.
Fix the probability generation rather than hiding or overriding outputs.
⸻
After every adjustment, verify that no subsystem silently reintroduces bias through normalization, smoothing, blending, calibration, or post-processing.
The final PMF should emerge directly from the simulation.
Do not artificially pull probabilities toward Overs, Unders, favorites, underdogs, close matches, or blowouts.
⸻
The rebuild is complete only when:
The guiding principle is simple: fix the probability engine, not the symptoms. If one generic improvement changes the shape of the PMF correctly, every downstream market should improve automatically.
════════════════════════════════════════
🎾 TENNISLOCKS 🎾
════════════════════════════════════════
🎯 WTA 250 Iasi | Clay, Outdoor | Best of 3 | Total line 20.5
Court speed index: 28 | Tour: WTA
────────────────────────────────────────
Tereza Valentova vs Aliaksandra sasnovich
────────────────────────────────────────
READING GUIDE: Forecast = model outcome | Fair price = model probability price | External EV = comparison with an entered sportsbook price | Pass = no priced edge
========================================
📊 MATCH WINNER
────────────────────────────────────────
Winner forecast: Tereza Valentova (62.4%) | fair ML -166
Moneyline: PASS at entered prices
Model fair ML: Tereza Valentova -166 / Aliaksandra sasnovich +166
════════════════════════════════════════
MATCH SUMMARY: Winner forecast: Tereza Valentova 62.7% | Match length: 2 sets
🎲 MATCH LENGTH:
Forecast: 2 sets / Under 2.5 (51.4%) | confidence very close
Model pick: Under 2.5 sets | probability 51.4% | model fair -106 | stability check passed
📊 Player Stats (Live):
========================================
🎯 TOTAL GAMES:
TOTAL FORECAST: OVER 20.5 (65.1%) | model fair odds -187
Mean 24.0 | median 23 | mode 19 (7.1%) | SD 5.7 | skew 0.20 | typical 80% range 17-31 games
Over probability composition: 16.6% from straight-set paths | 48.5% from three-set paths
Set-shape audit: 7-5 non-tiebreak 12.0% | 7-6 tiebreak 12.4%
Canonical hold inputs: Tereza Valentova 59.5% | Aliaksandra sasnovich 53.7%
########################################
🎯 PLAYER PROPS & PROJECTIONS 🎯
########################################
📊 Tereza Valentova - Projections:
Games Won: 12 projected on most likely 2-0 path (34.1%) | full PMF mean 12.7 | median 13 | mode 12
1st Set Games: OVER 5.5 (64%) | Fair -178
Sets Won: 1.45 projected | median 2 | mode 2 (62.4%) | P(0) 17.3% / P(1) 20.2% / P(2) 62.4%
Serve Games: 11.85 projected
Serve Points: 81.3 projected
Aces: UNDER 0.5 (71%) | Fair -240
Double Faults: UNDER 3.5 (54%) | Fair -118
Breaks Won: UNDER 5.5 (53%) | Fair -112
Break Points Created: OVER 10.5 (53%) | Fair -113
BP Conversion: 49% projected (5.5 / 11.3)
Opponent BP Save: 51% projected (5.8 saved / 11.3 faced)
Opponent return: opp return 28% | gate needs player sample/opponent sample
â ï¸ CAUTION: Props below computed from tour-average return rates due to insufficient player sample. Confidence is reduced.
📊 Aliaksandra sasnovich - Projections:
Games Won: 7 projected on most likely 2-0 path (34.1%) | full PMF mean 11.3 | median 12 | mode 12
1st Set Games: UNDER 5.5 (52%) | Fair -107
Sets Won: 1.04 projected | median 1 | mode 2 (37.6%) | P(0) 34.1% / P(1) 28.4% / P(2) 37.6%
Serve Games: 11.84 projected
Serve Points: 81.7 projected
Aces: OVER 0.5 (56%) | Fair -126
Double Faults: UNDER 4.5 (51%) | Fair -105
Breaks Won: OVER 4.5 (53%) | Fair -113
Break Points Created: OVER 9.5 (54%) | Fair -117
BP Conversion: 46% projected (4.8 / 10.4)
Opponent BP Save: 54% projected (5.6 saved / 10.4 faced)
Opponent return: opp return 29% | gate needs player sample/opponent sample
â ï¸ CAUTION: Props below computed from tour-average return rates due to insufficient player sample. Confidence is reduced.
🎲 Match-Level Context:
1st Set Total: UNDER 9.5 (51.1%)
Tiebreak forecast: NO TIEBREAK (72.2%)
Tiebreak O/U: UNDER 0.5 (72.2%)
pythonimport mpmath as mp p=0.624 f=lambda s: s*s*(3-2*s)-p s=mp.findroot(f,0.58) float(s), float(2*s*(1-s)) (0.5834412738376754, 0.48607510764069206)
Your diagnosis is substantially correct, and the uploaded code makes the failure mode more precise.
The relevant renderer:
jsfunction _gamesWonForecastPathValue_(engine,side){ ... var value=Number(st.mode); if(!Number.isFinite(value))value=Number(st.median); if(!Number.isFinite(value))value=Number(st.expected); ... }
It then prints that conditional path value first:
jspath.value.toFixed(0) + ' projected on most likely ' + path.score + ' path'
Only afterward does it report the full-PMF mean. Therefore:
That is exactly the “fixed the example rather than the engine” behavior you suspected.
The total-games engine directly consumes root.totalPmf, computes its statistics, and prices the line by summing its discrete mass. Its audit explicitly says:
decisionUsesMean:falsedecisionInputs:['CANONICAL_EXACT_TOTAL_GAMES_PMF','SPORTSBOOK_LINE']gatesApplied:0rescuesApplied:0otherEnginesDiagnosticOnly:trueSo the Over 20.5 probability of 65.1% is not a displayed-value hack or a post-PMF Over adjustment. It is genuinely what the current root distribution believes.
The conservation validator also checks that:
That means the downstream markets are now coherent with one another, but they can all be coherently wrong if the root is misshaped. (TennisLocksEngine_v416_SCORELINE_TRANSPARENCY_NO_SYNTHETIC_GW_LINES_20260716 (1).txt)
The strongest evidence does not point to a special 5-5, 6-5, or 6-6 transition bug.
It points to the way the BO3 engine couples winner probability to match length.
The BO3 engine:
jsvar targetP3 = winnerAlignedP3;
Although it computes a separate recent-history transition estimate, that estimate is diagnostic; it is not used as the live P3 target.
That creates an almost deterministic relationship between match-win probability and deciding-set probability.
For a stationary, approximately independent BO3 set model:
A match-win probability of 62.4% implies approximately:
Then:
The engine reports 48.5% from three-set paths.
That near-exact match is the architectural fingerprint. The high total is primarily being produced because a 62.4% winner is automatically being represented as a player who wins each set roughly 58% of the time, which naturally produces a deciding set almost half the time.
In other words:
The engine does not independently model “how often A wins” and “how decisively A wins.”
It derives both from nearly the same stationary strength parameterization.
Using the displayed canonical holds:
and reproducing the legal alternating-server set transitions gives approximately:
The report shows:
So those long-set frequencies are not being introduced by a hidden clutch rule in this canonical branch. They are the exact consequences of the supplied hold probabilities and ordinary tennis scoring.
This does not prove that the hold inputs are correctly calibrated. It means that suppressing 7-5 or tiebreak outcomes afterward would be the wrong repair. The real questions are:
The canonical state-space also keeps recent terminal-score and tiebreak history diagnostic-only rather than applying it to the live distribution.
Your existing prompt is strong, but this code-specific section would prevent the AI from spending its effort on the wrong subsystem:
:::writing{variant="document" id="58317"} 13. BREAK THE ARTIFICIAL WINNER-PROBABILITY-TO-MATCH-LENGTH COUPLING
The current BO3 architecture appears to adjust the players’ hold probabilities until the legal state-space reproduces the externally blended match-winner probability. It then accepts the deciding-set probability generated by that same stationary state-space.
This can create an artificial near-one-to-one relationship between match-winner probability and deciding-set probability.
Do not assume that a 62% match winner must be represented as a player with approximately a 58% chance of winning every set.
Winning probability and winning decisiveness must be independently identifiable.
Specifically audit and remove any architecture equivalent to:
The canonical score model must jointly represent at least:
These quantities must be calibrated jointly on held-out matches. Match-winner probability alone is not sufficient to determine the score-distribution shape.
Use a generic, symmetric model rather than favorite/underdog rules. Suitable approaches include:
Do not set the final BO3 deciding-set target automatically equal to the structural P3 produced after winner alignment.
Log and expose the following before and after reconciliation:
The rebuild must demonstrate that two matches with the same match-winner probability can produce different deciding-set and game-margin distributions when their hold levels, hold gaps, volatility, and calibrated score-shape evidence differ.
Do not label the modal player-games result within the most-likely exact-score path as the player’s overall Games Won projection.
The primary projection must be an unconditional statistic from the complete canonical PMF, such as:
A conditional value may be displayed separately as:
“Games on most-likely exact-score path.”
It must not replace or visually outrank the unconditional projection.
For example:
Games Won projection: mean 12.7 | median 13 | mode 12
Most likely exact-score path: 2-0, with 12 games won
Do not describe 12 as the overall projection merely because it is the modal result inside the most likely match-score branch. :::
The current result is not primarily:
“The totals engine added an Over adjustment.”
It is closer to:
“The winner engine supplied 62.4%, the score engine translated that into an approximately 58.3% per-set winner, and the resulting stationary BO3 structure naturally generated about 48.6% deciding sets.”
Then v416 put 12 and 7 in the headline by selecting the modal game counts inside the most-likely 2-0 path, while the unchanged PMF continued to produce 12.7, 11.3, and 24.0 underneath.
So your central conclusion is right: the display was repaired, but the score-distribution architecture was not independently recalibrated. (TennisLocksEngine_v416_SCORELINE_TRANSPARENCY_NO_SYNTHETIC_GW_LINES_20260716 (1).txt)
I would make the instructions even stronger. Right now your AI has correctly identified where the PMF comes from, but it is still assuming:
“The PMF is canonical, therefore the PMF is probably correct.”
That assumption is exactly what is preventing the real fix.
The next audit should start from the opposite assumption:
The canonical PMF itself may be structurally miscalibrated. Canonical does not automatically mean accurate. Every upstream assumption that shapes the PMF must be treated as a potential source of systematic error.
Something like this:
⸻
TENNISLOCKS ENGINE-WIDE SCORE DISTRIBUTION & TOTAL GAMES REBUILD
The previous rebuild successfully unified the downstream markets. Winner, Games Won, Total Games, Sets Won, and the renderer now appear to consume the same canonical probability object.
However, canonical consistency does not guarantee predictive correctness.
The remaining evidence indicates that the root score-distribution itself may be structurally biased.
The objective of this rebuild is not to make Overs become Unders or vice versa.
The objective is to determine whether the canonical probability distribution itself is producing systematically unrealistic tennis scorelines, match lengths, and total games.
Treat every upstream assumption as potentially incorrect.
Do not assume the current PMF is properly calibrated simply because all downstream markets now agree with it.
⸻
ISSUE 1 - The canonical PMF itself must be audited
The current engine assumes that once the canonical PMF exists, all derived markets become correct.
That assumption is unsafe.
Every downstream market can be internally consistent while still being wrong if the root PMF is misshaped.
Audit the construction of the canonical PMF itself.
Determine exactly how probability mass is assigned to every legal BO3 and BO5 scoreline.
Investigate whether probability is being allocated realistically across:
The canonical PMF should be treated as a model requiring calibration, not as a ground truth.
⸻
ISSUE 2 - Audit the score-shape generator instead of the renderer
The renderer is no longer the problem.
Games Won…
Sets Won…
Total Games…
Winner…
all now appear to consume the same PMF.
Therefore the investigation must move upstream.
Audit every subsystem responsible for generating the score distribution before reconciliation.
Determine whether the PMF already contains structural bias before any market is derived.
⸻
ISSUE 3 - Break the automatic coupling between winner probability and score shape
The current architecture appears to solve player strength until the blended winner probability is reached.
It then accepts the resulting score distribution with minimal independent calibration.
This creates a hidden assumption:
“If Player A wins 62% of matches, then the score distribution is whatever naturally emerges from those solved hold probabilities.”
That assumption is not generally valid.
Two players can both have:
Winner probability = 62%
while producing very different score distributions.
Examples include:
Winner probability alone should never determine score shape.
Score shape must be calibrated independently.
⸻
ISSUE 4 - Investigate systematic three-set inflation
The current PMF appears capable of assigning excessive probability to deciding sets.
Do not assume the deciding-set frequency is correct because it is mathematically implied by the current hold model.
Instead investigate whether the simulator itself systematically overproduces:
Determine whether the transition model inherently favors longer matches.
Audit:
If any stage consistently inflates deciding sets, repair the generator rather than compensating afterward.
⸻
ISSUE 5 - Investigate long-set inflation
Do not focus only on deciding sets.
Also audit individual set lengths.
Determine whether the engine systematically overproduces:
Likewise determine whether it underproduces:
The objective is not to reduce long sets.
The objective is to ensure long-set frequency matches realistic tennis scoring.
⸻
ISSUE 6 - Audit hold calibration
The displayed hold percentages are only intermediate inputs.
They are not automatically correct.
Investigate:
Determine whether these inputs consistently compress player differences.
If hold percentages become too similar…
the engine naturally produces:
without any explicit “Over bias.”
⸻
ISSUE 7 - Audit latent match variance
The engine currently appears to use relatively stationary player strengths.
Real tennis is not stationary.
Performance fluctuates within matches.
Investigate whether the engine properly represents:
If match variance is oversimplified…
the PMF may become unrealistically concentrated around close matches.
⸻
ISSUE 8 - Audit score-margin calibration
Do not calibrate only winner probability.
Also calibrate:
expected game differential
loser games
winner games
straight-set margins
three-set margins
terminal score frequencies
set-score frequencies
service-break frequencies
The engine should reproduce realistic score margins…
not simply realistic winners.
⸻
ISSUE 9 - Investigate Total Games overproduction
Do not assume the current Over frequency is correct simply because it is generated by the PMF.
Audit whether the PMF systematically overproduces:
Likewise audit whether it underproduces:
Determine whether the PMF itself possesses a statistical bias toward longer matches.
If such a bias exists…
repair the score generator rather than adjusting Over probabilities afterward.
⸻
ISSUE 10 - Audit distribution shape instead of only averages
Means alone are insufficient.
For every generated PMF examine:
mode
median
mean
variance
skewness
kurtosis
tail probability
line crossing probability
exact-score frequency
A correct mean can still hide an incorrect distribution.
The engine must reproduce realistic distribution shapes rather than merely realistic averages.
⸻
ISSUE 11 - Eliminate hidden structural feedback loops
Investigate whether any subsystem unintentionally reinforces another.
Examples include:
winner calibration changing holds
holds changing score margins
score margins changing deciding-set probability
deciding-set probability changing totals
totals influencing props
props influencing confidence
confidence influencing reconciliation
No subsystem should recursively bias another unless that dependency is mathematically required.
⸻
ISSUE 12 - Verify no hardcoded display substitutions remain
Do not assume any displayed statistic represents the true PMF.
Trace every displayed value back to its originating probability object.
Verify that:
If any display value is conditional, clearly label it as conditional and never allow it to replace the unconditional projection.
⸻
ISSUE 13 - Replace assumption-driven calibration with evidence-driven calibration
Do not preserve existing formulas simply because they are mathematically elegant.
Every component should justify itself by improving predictive accuracy.
If a simpler model consistently predicts better than a more complex model, prefer the simpler model.
If historical evidence contradicts the current architecture, revise the architecture rather than forcing reality to fit the model.
⸻
FINAL OBJECTIVE
The goal is not to “make this one match predict Under 20.5.”
The goal is to rebuild the score-distribution engine so that, across all ATP, WTA, Challenger, and ITF matches:
all emerge naturally from a well-calibrated canonical probability distribution.
Do not hardcode corrections for this match or for favorites/underdogs. If the rebuilt engine is correct, this match’s forecast may still be Over 20.5, or it may become Under 20.5. Either outcome is acceptable only if it results from a genuinely improved and better-calibrated score-distribution model, not from display changes, manual offsets, heuristic overrides, or special-case logic.
Audit and fix and edit my script then send me the script
I audited and edited the uploaded v416 engine. (TennisLocksEngine_v416_SCORELINE_TRANSPARENCY_NO_SYNTHETIC_GW_LINES_20260716 (1).txt)
Validation completed:
This is an architectural correction, not proof of empirical calibration. Proper ATP/WTA/Challenger/ITF calibration still requires a held-out historical score dataset.
Download the complete v417 script
This is a much more meaningful architectural change than the earlier “display fixes.” The key claims, if implemented correctly, attack a genuine source of bias instead of hiding it.
That said, I still would not consider the Total Games engine “fixed” yet, because there’s another layer that can still distort totals even after decoupling winner probability from score shape.
The next instructions I’d give your AI would be:
⸻
ENGINE-WIDE TOTAL GAMES DISTRIBUTION AUDIT (POST-v417)
The previous rebuild successfully removed the artificial coupling between winner probability and score shape. That was an important architectural correction. However, this only eliminates one potential source of bias. It does not prove that the score-distribution model itself is correctly calibrated.
The next objective is to determine whether the canonical score generator is still systematically producing unrealistic Total Games distributions.
The Total Games PMF is ultimately the sum of thousands of legal score paths. A bias anywhere upstream can distort the final Over/Under probability even if the PMF is internally consistent.
Trace every stage contributing probability mass into the Total Games PMF, including:
For each stage, measure how much it changes the final Total Games distribution. No stage should silently inject extra mass toward longer or shorter matches without explicit mathematical justification.
⸻
Do not evaluate only:
Instead inspect the entire scoreline distribution.
Measure the frequency of every legal outcome:
for each set.
Determine whether the generator systematically:
A Total Games model can appear reasonable while still allocating probability incorrectly across scorelines.
⸻
Winner probability answers:
“Who is more likely to win?”
It does not answer:
“How competitive will the match be?”
Determine whether the engine independently models:
If competitiveness is derived only from player strength, then the engine will naturally converge toward overly similar match shapes.
Introduce an independently calibrated competitiveness component if one does not already exist.
⸻
Real matches exhibit performance variation that is not constant throughout the match.
Determine whether the simulator currently assumes stationary serving and returning ability.
If so, investigate whether introducing realistic latent performance variation improves scoreline realism without introducing artificial noise.
The objective is not randomness.
The objective is realistic dispersion.
⸻
Small errors in hold estimation compound dramatically over an entire match.
Investigate whether the estimation process compresses player holds toward the mean through:
If both players’ hold rates become artificially similar, the engine will naturally produce:
even without any explicit totals bias.
⸻
The PMF reconciliation process should preserve probability conservation.
Verify that reconciliation does not unintentionally:
Reconciliation should satisfy constraints while introducing the smallest mathematically necessary adjustment.
⸻
Treat Total Games as its own predictive problem.
Determine whether the PMF systematically:
Identify whether any bias is:
Repair the generator rather than applying offsets to Total Games probabilities.
⸻
Every rebuild must preserve:
Winner ↔ Exact Score
Exact Score ↔ Sets Won
Exact Score ↔ Player Games
Exact Score ↔ Total Games
Exact Score ↔ Match Length
Exact Score ↔ Tiebreak Frequency
No market should diverge from the canonical PMF.
⸻
Final objective
Do not optimize the engine so that this specific match becomes an Under or so that any particular forecast changes.
Instead, ensure that the canonical score generator produces realistic tennis score distributions across all tours, surfaces, and player strengths.
A forecast should change only because the underlying model better represents real tennis scoring, never because of manual adjustments, hardcoded rules, favorite/underdog logic, or display-layer corrections.
⸻
This follows naturally after the v417 changes. The winner/score-shape coupling was one likely source of distortion, but it’s not the only one. The remaining work is to verify that the score generator itself produces realistic scoreline distributions and Total Games probabilities, rather than assuming the PMF is correct simply because it is internally consistent.
I completed the post-v417 audit and produced v418.
The engine already contained a capped, row-preserving terminal-score posterior for:
6-0, 6-1, 6-2, 6-3, 6-4, 7-5, 7-6
But it was not affecting the live PMF.
setDist() accepted score-shape evidence, and TennisSetShapeModel.apply() calculated the posterior, but the live generator returned the original structural distribution. The posterior was diagnostic-only, despite other comments claiming it had already been applied upstream.
That was a genuine engine-path omission rather than a display problem.
All completed successfully:
With no terminal-score evidence, v418 produces exactly the same score and Total Games PMFs as v417. It therefore does not invent a new cold-start adjustment.
A controlled evidence test changed mean Total Games from 24.00 to 22.95, while match-win probability remained unchanged to numerical precision. That demonstrates the repair can alter score shape without altering winner strength; it is not evidence that this particular reduction is empirically optimal.
Latent within-match performance variation remains disabled because enabling it without held-out calibration would merely replace one assumption with an arbitrary noise coefficient.
Download TennisLocks v418 JavaScript
Download TennisLocks v418 text version
Download the validation report
This repairs another real architectural omission, but it still does not prove that the hold priors, posterior concentration, or tour-specific score frequencies are empirically calibrated. That final determination requires held-out historical score data.
This is actually a much stronger change than the previous ones. The AI is now fixing real architectural omissions instead of papering over contradictions.
But I still see one major area that I would push on before I’d consider the totals engine “finished.”
The AI keeps saying:
“This changes score shape.”
The question it hasn’t answered is:
“What determines score shape when there is little or no historical terminal-score evidence?”
That’s likely where your remaining prediction error lives.
The engine is still fundamentally driven by hold estimates, and everything else is now layered on top of those holds.
That means if the hold model itself is slightly wrong, the entire PMF remains slightly wrong.
I’d send something like this next:
⸻
TENNISLOCKS v419 - HOLD MODEL & SCORE GENERATOR AUDIT
The previous rebuilds successfully removed multiple architectural issues:
These are significant improvements.
However, the remaining source of predictive error is likely no longer the PMF architecture itself.
It is likely the inputs that generate the PMF.
Do not assume the current hold probabilities are sufficiently accurate simply because the downstream engine is now coherent.
⸻
ISSUE 1 - Audit the hold model itself
Every match begins with estimated serve-hold probabilities.
Those values determine:
Therefore, even a small systematic error in hold estimation will propagate through the entire engine.
Trace every adjustment applied before the final hold values are produced.
Audit:
Determine whether any stage consistently compresses player differences or systematically overstates or understates hold ability.
⸻
ISSUE 2 - Measure sensitivity
Treat the hold model as a sensitivity problem.
For representative matches, perturb each player’s finalized hold probability by small amounts (for example ±0.5%, ±1%, ±2%).
Measure how much each perturbation changes:
Identify which markets are overly sensitive to small hold errors.
If a tiny change produces a disproportionate shift in Total Games, investigate why.
⸻
ISSUE 3 - Audit cold-start behavior
The previous rebuild correctly leaves cold-start matches unchanged when no terminal-score evidence exists.
Now determine whether cold-start matches rely too heavily on generic priors.
Audit:
Ensure uncertainty is represented by wider distributions rather than by forcing the player toward tour averages.
⸻
ISSUE 4 - Audit uncertainty propagation
The engine currently appears to use point estimates for hold probabilities.
Investigate whether uncertainty in those estimates should also propagate through the simulation.
Instead of assuming:
Player A Hold = 63.4%
consider whether the engine should internally represent a distribution around that estimate when confidence is low.
Low-information matches should naturally produce wider score distributions than high-information matches.
Do not inject arbitrary randomness.
Use uncertainty only where supported by the quality of the underlying data.
⸻
ISSUE 5 - Audit calibration of break frequency
Hold percentages alone do not guarantee realistic break distributions.
Verify that the resulting break frequencies match realistic tennis scoring.
Investigate whether the engine systematically:
Break frequency should emerge naturally from calibrated hold estimates rather than requiring downstream correction.
⸻
ISSUE 6 - Audit score-shape priors
When terminal-score evidence is sparse, determine exactly which priors dominate the PMF.
Document the relative influence of:
The engine should never default to an implicit assumption that all matches have similar score shapes.
⸻
ISSUE 7 - Verify no hidden feedback remains
Trace the complete dependency graph.
Confirm that none of the following feed back into score generation after the PMF is built:
The PMF should only depend on match characteristics and calibrated model inputs.
⸻
ISSUE 8 - Improve realism without heuristics
Do not add favorite rules.
Do not add underdog rules.
Do not add Over corrections.
Do not add Under corrections.
Do not manually rebalance scorelines.
If improvements are necessary, they should arise from better estimation of the underlying match process rather than from downstream adjustments.
⸻
At this point, your AI has genuinely moved from fixing architecture to improving the statistical model, which is exactly where you want to be. If it successfully completes this stage, you’re no longer chasing display inconsistencies. You’re refining the quality of the simulation itself.
I completed the v419 Hold Model & Score Generator Audit.
Two actual hold-pipeline defects were corrected:
kmEloBridge() returns f_ij and f_ji, but v418 attempted to blend nonexistent fA and fB fields. v419 now damps the fields that are actually converted into hold probabilities.service games × 10. v419 now prefers direct service-point attempts, then observed points per service game, and uses the old multiplier only as a fallback.The larger architectural change is that hold uncertainty now reaches the canonical PMF:
BO3 now also records the requested ±0.5%, ±1%, and ±2% hold sensitivity tests for both players, including changes in winner probability, deciding sets, mean totals, line-crossing probability, tiebreaks, player games, and modal exact score.
The old winner-aligned hold solver is now disabled by default because it is not part of the live model. It remains available only as an optional research diagnostic.
Completed successfully:
In a synthetic exposure test, posterior hold uncertainty and its PMF impact decreased monotonically as service-point evidence increased. This validates the intended behavior, but it does not prove that the existing Efron–Morris variance constants or hold blend coefficients are empirically optimal.
Download TennisLocks v419 JavaScript
Download TennisLocks v419 text version
Download the v419 validation report
The remaining major task is empirical hold-model calibration on held-out ATP, WTA, Challenger, and ITF data. v419 now exposes the necessary stages and sensitivities so that calibration can target the actual source of error rather than adjusting Total Games afterward.
I’d make the prompt much more demanding. Right now the AI keeps saying “I fixed X” because you’re asking it to fix the engine. Instead, make it prove where the error comes from before it writes a single line of code.
⸻
TENNISLOCKS v418 ROOT-CAUSE INVESTIGATION (NO MORE ARCHITECTURAL PATCHES)
The engine has already undergone multiple architectural rebuilds (PMF unification, winner decoupling, score-shape routing, reconciliation fixes, Games Won fixes, etc.).
From this point forward do not assume another architecture bug exists.
Your job is to determine why the canonical model still generates Total Games distributions that appear too high despite internally consistent PMFs.
This is now a scientific investigation, not another rewrite.
⸻
RULE 1
Do NOT modify:
unless you first prove one of those components is mathematically wrong.
Do not “improve” anything because it feels suspicious.
Everything must be evidence-driven.
⸻
RULE 2
Treat the current PMF as innocent until proven guilty.
Assume:
Your task is discovering why those mathematically correct probabilities differ from real tennis.
⸻
PHASE 1
COMPLETE CONTRIBUTION ANALYSIS
Take one match.
Do not immediately output Over or Under.
Instead completely decompose the PMF.
Report exactly where every percentage point comes from.
For example
Over 20.5 = 60.5%
Break it into
2-set matches
Continue through
7-6 7-6
Then
3-set matches
etc.
Show
Probability
Average games
Contribution to Over
Contribution to Under
Contribution to Push
until
ALL probability sums equal exactly
100%.
Nothing hidden.
⸻
PHASE 2
IDENTIFY WHICH SCORELINES CREATE THE OVER
Do NOT simply say
41%
comes from 3 sets.
That isn’t enough.
Instead report
Example
Straight sets
6-4 6-4
adds
X%
to Over.
6-4 7-5
adds
Y%
to Over.
7-5 6-3
adds
Z%.
etc.
Then rank
Top 20 scorelines
by contribution.
This immediately identifies
whether
Over
comes from
Too many
6-4s
Too many
7-5s
Too many
7-6s
Too many
three setters
or something else.
⸻
PHASE 3
SENSITIVITY ANALYSIS
Now freeze everything.
Move only ONE parameter.
Example
Increase Player A hold
0.5%
Recalculate.
Undo.
Increase Player B hold
0.5%
Undo.
Increase break rate
Undo.
Increase volatility
Undo.
Increase latent variance
Undo.
Reduce deuce survival
Undo.
Reduce tiebreak entry
Undo.
Reduce hold uncertainty
Undo.
Reduce break conversion
Undo.
Reduce momentum influence
Undo.
Reduce fatigue influence
Undo.
Produce
partial derivatives.
Example
+1% hold
changes mean games
+0.42
+1 SD volatility
changes mean games
+0.71
+1% tiebreak frequency
changes Over
+2.4%
etc.
Now we know what actually drives totals.
⸻
PHASE 4
ELIMINATION TEST
Disable ONE subsystem at a time.
Do NOT delete code.
Temporarily bypass it.
Measure
Mean games
Median
Mode
Over %
Expected loser games
Expected favorite margin
Expected three-set %
Expected tiebreak %
Repeat for
Hold shrinkage
Hold uncertainty
Volatility
Momentum
Fatigue
Surface adjustment
Style adjustment
Serve model
Break model
Terminal-score posterior
Recent-score posterior
Historical score frequencies
Every independent adjustment.
This identifies the dominant contributor.
⸻
PHASE 5
HISTORICAL RESEARCH
After identifying the dominant subsystem,
STOP coding.
Research only that subsystem.
Examples
If
Hold estimation
is responsible
Research
peer-reviewed tennis hold estimation.
If
Volatility
is responsible
Research
within-match performance variance.
If
Break conversion
is responsible
Research
break point conversion models.
If
Score distribution
is responsible
Research
empirical ATP/WTA set-score frequencies.
Do not invent constants.
Use published work.
⸻
PHASE 6
HISTORICAL VALIDATION
For ATP
WTA
Challenger
ITF
compare
Generated frequencies
vs
Historical frequencies.
At minimum compare
6-0
6-1
6-2
6-3
6-4
7-5
7-6
Three-set frequency
Tiebreak frequency
Mean loser games
Mean favorite games
Mean total games
Game margin
Straight-set total
Three-set total
RMSE
MAE
KL divergence
Jensen-Shannon divergence
Brier score
Log loss
for every candidate calibration.
⸻
PHASE 7
ONLY THEN MODIFY THE MODEL
Only modify parameters whose adjustment measurably reduces historical error.
Reject every change that
Improves one statistic
while making others worse.
⸻
HARD REQUIREMENTS
Do NOT
Every improvement must emerge naturally from the underlying tennis model.
⸻
REQUIRED DELIVERABLE
Do not simply say
“I fixed totals.”
Instead produce a report like:
If you cannot prove the source of the error with quantitative evidence, do not modify the engine.
⸻
This prompt changes the AI’s job from “keep tweaking totals until they look right” into “perform a forensic audit of the entire score-generation pipeline and prove exactly which model assumption is inflating totals.” That’s a much stronger direction because it forces evidence before edits, reducing the chance of another round of patches that simply move the output without addressing the underlying cause.
I followed the evidence-first instruction and did not modify the engine.
The Iasi forecast is reproducible using only:
That reconstruction produces:
| Metric | Result |
|---|---|
| Mean Total Games | 24.0002 |
| Over 20.5 | 65.14% |
| Three sets | 48.61% |
| Any tiebreak | 27.84% |
| Player A games | 12.7003 |
| Player B games | 11.2999 |
The complete enumeration contains 1,470 legal ordered score paths, sums to exactly 100%, and reproduces the engine’s Total Games PMF to numerical precision.
The largest straight-set Over contributors are combinations such as:
6-4 7-57-5 6-46-4 7-67-6 6-46-3 7-56-3 7-6So the excess is not merely “too many three setters.” The engine also concentrates too much within-set probability around competitive terminal scores.
Reconciliation is effectively innocent:
Other findings:
Therefore, the high distribution already exists in the stationary hold-driven legal state space before reconciliation.
I compared the reconstruction with completed WTA clay matches from 2021–2025 using Jeff Sackmann’s public match-result data. The archive identifies Sackmann’s ATP and WTA repositories as its original sources and includes match data through 2026. (GitHub)
A 108-match comparison cohort with similar rolling pre-match clay hold estimates averaged:
| Metric | Historical cohort | Model |
|---|---|---|
| Three sets | 37.04% | 48.61% |
| Mean games | 21.75 | 24.00 |
| Over 20.5 | 50.00% | 65.14% |
| 6-0 | 6.64% | 3.12% |
| 6-2 | 21.88% | 16.47% |
| 6-4 | 18.75% | 24.50% |
| 7-6 | 7.81% | 12.39% |
The model is shifting mass away from decisive scores and toward close scores.
The match-level differences were statistically meaningful in this cohort:
p = 0.0162p = 0.00161p = 7.15e-05Diagnostic counterfactuals—not production recommendations—showed:
| Counterfactual | Mean change | Over change |
|---|---|---|
| Historical P3 only | −1.123 games | −7.78pp |
| Historical terminal set shape only | −1.188 games | −5.42pp |
| Both | −2.250 games | −14.33pp |
The remaining inflation appears approximately split between:
Published research also supports testing departures from stationary, independent-set assumptions: Depken, Gandar, and Shapiro report strategic and psychological set-level effects using 66,262 professional BO3 matches, while Cui and colleagues document set-to-set performance variation. Neither paper supplies a universal coefficient that should be inserted without held-out validation. (IDEAS/RePEc)
No v420 was created.
The evidence identifies the likely statistical failure, but it does not yet justify choosing a replacement model or coefficient. Production changes should wait for time-split held-out validation across ATP, WTA, Challenger, and ITF.
Download the full investigation report
Download the complete 1,470-path decomposition
TENNISLOCKS v420 ENGINE-WIDE ROOT CAUSE AUDIT & CANONICAL REBUILD
You are performing a complete engineering audit of the TennisLocks engine.
Assume there are still undiscovered architectural, mathematical, statistical, calibration, dependency, and reconciliation defects.
Do not assume any subsystem is correct simply because previous validations passed.
Previous versions fixed several architectural issues, but live outputs still demonstrate statistically unrealistic Total Games, Games Won, terminal score distributions, and deciding-set frequencies.
Your objective is to locate every remaining root cause, repair it, recalibrate it, validate it, and replace any flawed architecture.
⸻
PRIMARY OBJECTIVE
Treat every subsystem as potentially incorrect.
Audit every dependency from raw inputs through rendered outputs.
Do not stop after finding one issue.
Continue searching until every market is internally consistent, historically calibrated, and statistically validated.
If fixing one subsystem exposes another weakness, continue repairing until no additional structural issues remain.
⸻
PHASE 1
Complete dependency audit
Trace every dependency for:
For every displayed probability identify:
Produce a dependency tree.
Repair every inconsistency discovered.
⸻
PHASE 2
Remove hidden coupling
Search for any place where one market unintentionally controls another.
Examples include:
Winner influencing score shape
Winner influencing deciding sets
Winner influencing total games
Winner influencing player games
Hold calibration influencing unrelated props
Match length influencing exact score
Exact score influencing hold calibration
Any recursive dependency
Remove unintended coupling.
Only mathematically necessary dependencies should remain.
⸻
PHASE 3
Audit every probability distribution
Verify:
Normalization
Conservation
Expectation
Variance
Median
Mode
Skew
Tail behavior
Terminal score frequencies
Conditional probabilities
Joint probabilities
Marginal probabilities
If any distribution is mathematically inconsistent, rebuild it.
⸻
PHASE 4
Historical calibration rebuild
Do not simply compare one match.
Audit thousands of historical ATP, WTA, Challenger and ITF matches.
Calibrate:
Winner
2-set frequency
3-set frequency
Mean games
Median games
Games won
Loser games
Winner games
6-0
6-1
6-2
6-3
6-4
7-5
7-6
Tiebreak frequency
Break frequency
Service holds
Return games
Calibrate the model until historical error is minimized simultaneously.
⸻
PHASE 5
Statistical error localization
Measure contribution of every subsystem to Total Games error.
Examples
Hold estimation
Hold shrinkage
Surface adjustment
Fatigue adjustment
Momentum adjustment
Volatility adjustment
Score generator
State space
Markov model
Monte Carlo
Reconciliation
Posterior calibration
Score shape calibration
Latent variance
Break clustering
Terminal score calibration
Quantify how much each subsystem contributes.
Do not guess.
⸻
PHASE 6
Automatic replacement
If any subsystem consistently produces historical error:
Replace it.
Do not preserve architecture simply because it already exists.
If a simpler model performs better,
replace the complex one.
If a different statistical approach performs better,
replace the existing implementation.
⸻
PHASE 7
Calibration optimizer
Build automatic parameter optimization.
Optimize against held-out historical data.
Tune:
hold shrinkage
volatility
latent variance
break clustering
terminal score frequencies
score shape
deciding-set frequency
Do not use manually chosen constants if they can be learned.
⸻
PHASE 8
Cross-market validation
After every rebuild verify:
Winner
↓
Exact score
↓
Match length
↓
Games won
↓
Total games
↓
Props
Every market must agree.
No contradictory outputs.
No impossible combinations.
No statistical drift.
⸻
PHASE 9
Stress testing
Run calibration against:
Heavy favorites
Coin flips
Clay
Hard
Grass
Indoor
ATP
WTA
Challenger
ITF
Qualifiers
Retired players
High hold
Low hold
Extreme hold gaps
Very low hold gaps
Every configuration must remain calibrated.
⸻
PHASE 10
Automatic anomaly detection
Search the engine for:
legacy code
dead code
duplicate calculations
unused variables
hidden overrides
double calibration
stale constants
incorrect assumptions
floating-point accumulation
recursive feedback
incorrect normalization
overfitting
underfitting
redundant reconciliation
Replace or remove everything that contributes measurable error.
⸻
FINAL REQUIREMENT
Do not stop after fixing one issue.
Continue auditing until every subsystem has been inspected.
If repairing one subsystem reveals another statistical defect, continue repairing.
The final engine should be treated as a complete rebuild of the statistical pipeline rather than a sequence of patches.
Validate every repair with historical testing before accepting it.
Do not preserve existing architecture unless it demonstrably outperforms alternatives.
This is an evidence-backed rebuild rather than another display or totals adjustment.
6-0 through 7-6.The historical files came from the Jeff Sackmann ATP and WTA repositories and their archive mirror. (GitHub)
Across the accepted cells:
| Metric | Baseline | v420 |
|---|---|---|
| P3 log loss | 0.6819 | 0.6484 |
| Terminal-score log loss | 1.8221 | 1.8071 |
| Total Games MAE | 5.4472 | 4.9492 |
| Over 20.5 Brier | 0.2708 | 0.2504 |
| Exact-total log score | 3.1146 | 3.0385 |
| Total-distribution CRPS | 3.5140 | 3.2600 |
| Metric | v419 | v420 |
|---|---|---|
| Valentova winner probability | 62.40% | 62.40% |
| Mean Total Games | 24.000 | 21.809 |
| Over 20.5 | 65.14% | 50.68% |
| Three sets | 48.61% | 36.68% |
| Any tiebreak | 27.84% | 21.47% |
| Valentova games | 12.700 | 11.669 |
| Sasnovich games | 11.300 | 10.140 |
The winner probability is unchanged. The movement comes from historically calibrated match-length and terminal-score distributions.
The reference PMF still contains 1,470 legal ordered paths, sums to one, and reproduces the canonical Total Games PMF with total variation below 2e-16.
6.67e-163.20e-141e-14ITF score calibration and BO5 recalibration remain intentionally unchanged because the available data did not support defensible hold-conditioned replacements. The report documents these limits rather than presenting them as solved. The original engine lineage audited for this work is available in the uploaded source. (TennisLocksEngine_v416_SCORELINE_TRANSPARENCY_NO_SYNTHETIC_GW_LINES_20260716 (1).txt)
Download the v420 JavaScript engine
Download the v420 text version
Read the complete engineering audit
Download the validation report
Here’s a stronger version. It doesn’t just tell the AI to “look for problems.” It instructs it to actively break its own work, replace anything that doesn’t survive, and continue until the engine is statistically defensible.
⸻
TENNISLOCKS v421 RED TEAM AUDIT, RECALIBRATION & ENGINE REBUILD
Assume v420 is still wrong.
Forget every previous conclusion.
Treat every calibration, coefficient, learned parameter, lookup table, bucket, multiplier, shrinkage function, reconciliation routine, historical split, and validation result as potentially incorrect.
Your objective is not to defend v420.
Your objective is to disprove it.
Assume hidden statistical errors still exist until proven otherwise.
If you discover a better model, calibration, statistical approach, or architecture, replace the existing implementation.
Do not preserve code because it already exists.
⸻
PHASE 1
Audit everything added in v420
Audit every new component individually.
Including but not limited to:
Assume every one of these may still contain statistical defects.
⸻
PHASE 2
Attempt to falsify every calibration
For every calibration ask:
Can I prove this is wrong?
Attempt to falsify using
Historical replay
Cross validation
Leave-one-year-out validation
Leave-one-tour-out validation
Leave-one-surface-out validation
Bootstrap testing
Sensitivity analysis
Monte Carlo perturbation
Coefficient perturbation
Random initialization
Adversarial stress testing
If the calibration cannot survive these tests,
remove or rebuild it.
⸻
PHASE 3
Search specifically for
Data leakage
Future information leakage
Target leakage
Overfitting
Bucket sparsity
Small sample bias
Selection bias
Survivorship bias
Calibration drift
Distribution drift
Historical era bias
Tour imbalance
Surface imbalance
Tournament imbalance
Hold estimation bias
Player-strength bias
Favorite bias
Underdog bias
Ranking bias
Elo interaction bias
Momentum interaction bias
Unintended nonlinear interactions
Hidden hardcoded adjustments
Implicit thresholds
Legacy constants
Duplicated corrections
Double calibration
Recursive calibration
Coefficient cancellation
Compensating errors
Unsupported extrapolation
Silent clipping
Artificial probability compression
Artificial probability expansion
Numerical instability
Floating-point accumulation
Any discovered issue must be repaired.
Continue searching after every repair.
⸻
PHASE 4
Remove compensating errors
Search for situations where
Subsystem A is wrong
Subsystem B partially corrects A
making validation appear acceptable.
Eliminate the root cause.
Never preserve compensating errors.
⸻
PHASE 5
Independent recalibration
If any calibration fails,
rebuild it completely.
Do not tweak.
Retrain from historical data.
Generate entirely new parameters.
Replace the old calibration.
Repeat validation.
⸻
PHASE 6
Global optimization
Do not optimize one metric.
Simultaneously optimize
Winner probability
Moneyline calibration
Sets
Exact score
Games won
Total games
Three-set frequency
Terminal score frequencies
Tiebreak frequency
Break frequency
Hold calibration
Serve calibration
Prop calibration
CRPS
Brier
Log loss
Calibration slope
Calibration intercept
Expected calibration error
Reliability curves
The objective is minimum overall historical prediction error.
Not one improved statistic.
⸻
PHASE 7
Cross-market conservation
After every rebuild verify
Winner
↓
Exact Score
↓
Match Length
↓
Player Games
↓
Total Games
↓
Props
No market may contradict another.
Every displayed probability must originate from the canonical PMF.
⸻
PHASE 8
Calibration robustness
A calibration is accepted only if
it improves
ATP
WTA
Challenger
ITF
Clay
Hard
Grass
Indoor
Favorites
Underdogs
Coin flips
Heavy favorites
Low hold
High hold
Large hold gaps
Small hold gaps
Different seasons
If improvement occurs only in one segment,
reject the calibration.
⸻
PHASE 9
Replace weak models
If another statistical model performs better,
replace the existing implementation.
Examples include
Bayesian hierarchical calibration
Generalized additive models
Gradient boosting
Monotonic regression
Isotonic calibration
Spline calibration
Mixture models
Latent performance models
Hierarchical score-shape models
Do not preserve the existing architecture simply because it already works.
⸻
PHASE 10
Adversarial engine testing
Generate thousands of synthetic matches.
Generate historical replay tests.
Generate extreme edge cases.
Attempt to make the engine fail.
Examples
95% favorites
50/50 matches
Very low holds
Very high holds
Huge hold gaps
Tiny hold gaps
Very slow clay
Fast grass
Indoor hard
If any instability appears,
repair it.
⸻
PHASE 11
Quantify every retained calibration
For every coefficient,
table,
multiplier,
or learned parameter,
produce evidence showing
Why it exists
Historical support
Sample size
Confidence interval
Validation improvement
Sensitivity
Generalization performance
If this evidence cannot be produced,
remove that calibration.
⸻
PHASE 12
Final rebuild
After completing the audit,
perform another complete dependency audit.
Do not assume earlier repairs remain valid.
Search for newly introduced interactions.
Repair them.
Repeat until no statistically significant defects remain.
⸻
FINAL REQUIREMENT
Your task is not to produce v421.
Your task is to determine whether v420 deserves to exist.
If any part of v420 cannot survive an adversarial statistical audit, replace it.
Continue auditing, recalibrating, rebuilding, and validating until the engine is supported by evidence rather than previous assumptions.
Do not stop after finding one issue.
Assume another issue exists until exhaustive analysis demonstrates otherwise.
v420 did not survive unchanged. Its population calibrations generalized, but several live layers failed independent testing and were replaced or removed.
Player-specific deciding-set history overlay
On the independent 2026 replay of 6,935 matches:
| P3 model | Log loss | ECE |
|---|---|---|
| Structural hold model | 0.681520 | 0.125210 |
| Population calibration | 0.649562 | 0.008016 |
| Population + player history | 0.655194 | 0.030858 |
The history overlay matched the aggregate P3 rate more closely but destroyed calibration discrimination: slope fell from approximately 1.011 to 0.223. This was a compensating error, not a genuine improvement.
Player-specific terminal-score posterior
Across 16,327 completed sets in the 2026 replay:
| Terminal-score model | Log loss |
|---|---|
| Structural | 1.826410 |
| Population calibration | 1.808545 |
| Population + player history | 1.843131 |
The player-history layer worsened results in every evaluated year and every evaluated tour/surface cell. Both recent-history systems remain available as diagnostics but can no longer alter the canonical PMF.
Hold-uncertainty mathematics
v419/v420 calculated posterior variance using service-point-rate parameters, then added the resulting standard deviation directly to game-hold probability. Those are different probability scales.
v421 now:
The improvement is small but consistent:
| Model | Total MAE | Exact-total log score | CRPS |
|---|---|---|---|
| Point holds | 5.012015 | 3.009826 | 3.275398 |
| Old hold-space uncertainty | 5.011652 | 3.009723 | 3.275280 |
| Correct SPW-space uncertainty | 5.010234 | 3.009430 | 3.274855 |
The held-out population P3 calibration improves every evaluated year:
| Year | Structural | Retained calibration |
|---|---|---|
| 2024 | 0.682223 | 0.649436 |
| 2025 | 0.682410 | 0.649672 |
| 2026 | 0.681520 | 0.649562 |
The population terminal-score calibration also improves 2024, 2025, and 2026.
A competing global logistic model produced a tiny aggregate 2026 improvement of approximately 0.00017 in P3 log loss, but failed to improve consistently across tours, surfaces, and hold-gap segments. It was therefore rejected rather than substituted on the strength of a negligible overall gain.
Three v420 service fallback changes failed the partial-2026 replay and were reverted:
0.03711 → 0.0300.02829 → 0.0340.05044 → 0.049The historical replay used chronologically ordered public ATP and WTA match-result data compiled by Jeff Sackmann and its archival mirror. (GitHub)
An extreme recent-history input moved v420’s deciding-set probability by 23.69 percentage points and changed its Total Games PMF by 0.241 TV distance. In v421, the same input changes both by exactly zero because the failed posterior is diagnostic-only.
Validation results:
4.44e-164.26e-14The original engine audited in this work is the uploaded TennisLocks source. (TennisLocksEngine_v416_SCORELINE_TRANSPARENCY_NO_SYNTHETIC_GW_LINES_20260716 (1).txt)
Download the v421 JavaScript engine
Download the v421 text version
Read the complete red-team audit
Download the validation report
Download the complete reproducible audit package
v421 is better supported than v420, but it is not presented as proof that every possible defect has been eliminated. ITF hold-conditioned calibration, indoor cells, and new BO5 calibration remain unresolved where the available data did not justify a replacement.