From 0b12137376929d96da81cd088d3d288cb1eec32d Mon Sep 17 00:00:00 2001 From: Scott Gasch Date: Sat, 29 Aug 2026 15:03:54 -0700 Subject: Dynamic move ordering overhaul: continuation-history, evidence-gated countermove promotion, retired hung-piece-escape and NumLeftoverMovesToSelect. Full session was built on a "measure the pick, not the game" methodology: aggregate solve counts on curated suites are too noisy to tune move-ordering knobs against, so most decisions here came from per-move fail-high/alpha-raise rates at much larger sample sizes (leftover FH% instrumentation, a zero- selection-budget diagnostic that isolates a single best-of-remaining pick, and evidence-bucket calibration), not solve-count deltas alone. See CLAUDE.md's "Dynamic move ordering experiments" section for the reusable methodology and generate.c's _ScoreAllMoves comment for the resulting ordering hierarchy. Changes: - Added g_ContinuationHistory: same growth/decay math as the existing g_HistoryCounters butterfly table, additionally keyed by the previous move, so its magnitude is self-calibrated rather than a hand-picked constant. Flat, sufficient response across a 256x scale sweep. - Countermove-table matches now get a real GOOD_MOVE-tier promotion (previously the table was write-only, tracked for stats but never read for ordering), but only when the match's own accumulated history+continuation evidence clears COUNTERMOVE_EVIDENCE_THRESHOLD (10,000) -- a raw match with no track record was shown to perform identically to an ordinary leftover (~0.6-0.85% FH), so promoting on match alone would have repeated hung-piece-escape's mistake below. - Retired hung-piece-escape's unconditional GOOD_MOVE-tier promotion. Evidence-calibration showed the overwhelming majority of triggers (a zero-evidence population 250-1000x larger than countermove's) performed at the plain-leftover baseline -- the promotion was mostly free tier- escape treatment for moves that hadn't earned it. Replaced with FLEE_BONUS, a flat same-tier nudge inside SelectBestWithHistory (never escapes GOOD_MOVE/leftover classification, unlike a generation-time promotion) at the magnitude found to plateau a same-tier-nudge sweep. - Retired NumLeftoverMovesToSelect (the depth-indexed budget on how many leftover moves got a full selection scan before falling back to unsorted order). search.c's main move loop now always fully selects -- the leftover pool was shown to contain real, findable signal a bailout budget was discarding for a node-count savings that didn't hold up net- net once measured by solve counts and fail-high rates rather than raw node counts (noisy on small suites independent of this change). - Collapsed leftover-move instrumentation from sorted/raw pairs down to a single set now that "raw" (unsorted fallback) is structurally impossible; kept the countermove evidence-bucket calibration counters (ongoing check that COUNTERMOVE_EVIDENCE_THRESHOLD stays well- calibrated); removed the contested-node A/B harness and hung-piece evidence calibration now that the decisions they were built to inform are made. Net effect on the three curated suites (sd 10): solve counts wash (tied, +1, -1 across ringers/confident/hard), leftover fail-high rate improved consistently on all three (the intended, directly-measured target of this work). Not yet validated beyond sd 10 -- an sn-based run or eval_tune/match_play.py head-to-head gate is the natural next check before leaning on this as a proven strength gain rather than a directionally- sound, sd-10-clean change. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_014XePz6Sk4qQsTaP2jVJWJu --- src/searchsup.c | 53 ----------------------------------------------------- 1 file changed, 53 deletions(-) (limited to 'src/searchsup.c') diff --git a/src/searchsup.c b/src/searchsup.c index 54e1814..72667c8 100644 --- a/src/searchsup.c +++ b/src/searchsup.c @@ -24,59 +24,6 @@ extern SCORE g_iRootScore[2]; extern ULONG g_uHardExtendLimit; extern ULONG g_uIterateDepth; -ULONG -NumLeftoverMovesToSelect(IN SEARCHER_THREAD_CONTEXT *ctx, IN ULONG uDepth) -/** - -Routine description: - - How many "leftover" (below GOOD_MOVE -- see search.c's - TRY_GENERATED_MOVES gate) moves are worth a full SelectBestWithHistory - scan before we give up and just take the remainder in whatever order - they're sitting in. Only ever consulted once every high-performer - move (winning/even capture, killer, killer-mate) has already been - exhausted -- this never limits how many of *those* get selected, - only how much further care to spend on the ordinary/leftover tail. - - Replaces the old g_uSearchSortLimits[ply], indexed by distance from - the root -- a poor proxy for what actually matters here, which is - how large the remaining subtree below this node is (distance from - root only correlates with that when total search depth is roughly - fixed; it says nothing once extensions/reductions/iterative-deepening - are in play). uDepth (remaining depth, in ONE_PLY units, possibly - fractional) is the more principled signal: a bigger remaining - subtree makes the cost of a few extra O(n) selection scans more - worth paying to avoid a bad early choice cascading into extra - full-width re-searches. - - STARTING POINT, NOT YET VALIDATED under this new meaning: this - reuses the previous table's six numbers verbatim, just reindexed by - plies of *remaining* depth instead of *distance from root* -- same - overall shape (more care with more depth left), same specific - values, carried over only because they're a known, testable - starting point, not because they were ever confirmed correct here. - -Parameters: - - IN SEARCHER_THREAD_CONTEXT *ctx, - IN ULONG uDepth - -Return value: - - ULONG - -**/ -{ - static const ULONG _uLimits[] = { 8, 9, 11, 13, 15, 17 }; - ULONG uPlies = uDepth / ONE_PLY; - - if (uPlies >= ARRAY_LENGTH(_uLimits)) - { - uPlies = ARRAY_LENGTH(_uLimits) - 1; - } - return(_uLimits[uPlies]); -} - void UpdatePV(SEARCHER_THREAD_CONTEXT *ctx, MOVE mv) /** -- cgit v1.3