summaryrefslogtreecommitdiff
path: root/src/board_representation/MIGRATION.md
diff options
context:
space:
mode:
Diffstat (limited to 'src/board_representation/MIGRATION.md')
-rw-r--r--src/board_representation/MIGRATION.md89
1 files changed, 62 insertions, 27 deletions
diff --git a/src/board_representation/MIGRATION.md b/src/board_representation/MIGRATION.md
index 57ca7f9..313c74d 100644
--- a/src/board_representation/MIGRATION.md
+++ b/src/board_representation/MIGRATION.md
@@ -15,14 +15,15 @@ inline as a note rather than deleted, so a future session designing
the separate migration doesn't have to rediscover it from git history.
**Status: sections 1-3 (`bbPieces`/`bbPawns`-backed
-`_WhoAttacksSquareBB`/`_GetAttacksBB`) and section 6 (the
-`GETATTACKS_BITBOARD` toggle, `_GetAttacksBB` now reachable from real
-search when enabled) implemented, verified, and committed. Section 4
-items 1-3 done (folded into sections 3/6's own testing); items 4-5 and
-sections 5's second half/7 -- the whole-engine curated-suite and
-`match_play.py` gates -- not started, now genuinely unblocked since
-section 6 exists. Section 8 not started.**
-See the per-section status notes below for specifics.
+`_WhoAttacksSquareBB`/`_GetAttacksBB`), section 4 items 1-3 and 5, and
+section 6 (the `GETATTACKS_BITBOARD` toggle) all implemented, verified,
+and committed -- including a whole-engine `sd10` check across all three
+curated suites showing zero solve-count regression and +8.38%
+aggregate NPS. Only `match_play.py` (section 4 item 6 / section 7)
+remains unmet before full retirement of the old mailbox
+implementation, deliberately deferred for now. Section 5's whole-engine
+item and section 7's remaining items are otherwise done; section 8 not
+started.** See the per-section status notes below for specifics.
Context: a prior pass of this session added occupancy-bitboard-driven
mobility/attack-presence bitboards (`bbPawnAttacks`, `bbMinorAttacks`,
@@ -467,16 +468,45 @@ sequence.
directly (`generate.c`'s 5 call sites, `searchsup.c:365`), so even
value-for-value-identical results can shift which move gets tried
first or which lines get pruned/extended, changing node counts and
- occasionally search results at the margins. Run all three curated
- suites (`ecm_ringers`, `ecm_confident_quick`, `ecm_hard_quick`) at
- `sd10` against `head_reference/`, and don't treat a clean
- `TestGetAttacks` unit-level pass as sufficient sign-off by itself.
- Not started -- requires section 6's toggle to exist first, so there's
- something to flip on for the suite runs.
+ occasionally search results at the margins. **DONE** -- run all
+ three curated suites at `sd10`, but *not* against the checked-in
+ `head_reference/` binary: that binary predates this branch's
+ sections 1-6 by a dozen-plus unrelated commits, so diffing against
+ it would conflate this change with everything else on the branch.
+ Instead, isolated the one variable that matters: same commit, built
+ twice (`gmake clean && gmake -j5 GENETIC=1 PERF_COUNTERS=1 MP=1
+ SIXTYFOUR=1` with and without `GETATTACKS_BITBOARD=1`), same
+ `--cpus 1 --hash 256m`, `sd 10`, `book name /nonexistent.book.bin`
+ invocation as `head_reference/`'s own convention. Results:
+
+ | suite | solves (asm -> BB) | nodes (asm -> BB) | NPS (asm -> BB) |
+ |---|---|---|---|
+ | `ecm_ringers` (11) | 10 -> 10 | 37,611,927 -> 37,909,197 (+0.79%) | 1,178,986 -> 1,270,436 (+7.76%) |
+ | `ecm_confident_quick` (90) | 83 -> 83 | 429,387,924 -> 430,930,314 (+0.36%) | 1,201,203 -> 1,300,133 (+8.24%) |
+ | `ecm_hard_quick` (90) | 25 -> 25 | 459,712,049 -> 465,403,216 (+1.24%) | 1,131,433 -> 1,228,991 (+8.63%) |
+
+ **Solve counts bit-identical on all three suites -- zero
+ regression.** Node counts up slightly (+0.36% to +1.24%), exactly
+ the "value-identical but not order-identical" effect warned about
+ above -- `_GetAttacksBB` returns the same attacker *set* as asm
+ `GetAttacks` but not necessarily the same order, so `SEE()`/move-
+ ordering tie-breaking can shift which lines get searched first at
+ the margins. Didn't cost a single solve here. Aggregate (total
+ nodes / total script time across all three suites, not an
+ average-of-averages): baseline 926,711,900 nodes / 795.7s =
+ 1,164,842 nps; bitboard 934,242,727 nodes / 740.0s = 1,262,490 nps
+ -- **+8.38% overall wall-clock NPS**, comfortably absorbing the
+ extra node count. Whole-engine result confirms section 3's isolated
+ cycles/call benchmark wasn't an artifact of measuring the wrong
+ thing -- the speedup shows up in real search throughput, not just
+ in the isolated per-call harness.
6. **`match_play.py` gate** (`LOWER95 >= 0.5`) before calling this done
-- same reasoning: this is closer to a search-behavior change than a
pure eval-magnitude change, so the existing eval-tuning gate criteria
- apply. Not started, same section-6 dependency as above.
+ apply. Not started yet, deliberately deferred -- item 5's zero solve
+ regression + throughput win is convincing enough on its own for now;
+ this gate is still required before full retirement (section 7), just
+ not being chased immediately.
## 5. Microbenchmarking
@@ -498,14 +528,17 @@ visible in isolation. Two tiers, both required before wiring in:
asm opening, 0.55x middlegame, 0.89x endgame). Consistent win across
the spectrum, no red flag.
2. **Whole-engine cycles-per-node / NPS**, not just the isolated call:
- `sd`-fixed-depth comparison against `head_reference` on the three
- curated suites, using the engine's own self-reported `Searched for N
- seconds, M nodes` (not `ps` sampling -- noisy on this shared box per
- earlier sessions). This is the check that would have caught the Eval
- regression earlier if it had been run before the reader-migration
- work proceeded -- don't skip straight from isolated-cycles-looks-good
- to production wiring again. Not started -- requires section 6's
- toggle to exist first.
+ `sd`-fixed-depth comparison on the three curated suites, using the
+ engine's own self-reported node counts/NPS (not `ps` sampling --
+ noisy on this shared box per earlier sessions). This is the check
+ that would have caught the Eval regression earlier if it had been
+ run before the reader-migration work proceeded -- don't skip
+ straight from isolated-cycles-looks-good to production wiring
+ again. **DONE**, folded into section 4 item 5's same-commit A/B run
+ (see there for the numbers and why `head_reference/` itself wasn't
+ the comparison point) -- +8.38% aggregate wall-clock NPS across all
+ three suites, confirming the isolated-cycles-per-call win (section
+ 3) actually shows up end to end, not just in the harness.
## 6. Dual-support / toggle strategy -- DONE
@@ -562,10 +595,12 @@ Only delete `SlowGetAttacks`/asm `GetAttacks` after **all** of:
- Isolated cycles/call shows a consistent win across the piece-density
spectrum (not just one favorable position). **DONE.**
- Whole-engine `sd10` on all three curated suites shows no solve-count
- regression vs. `head_reference`. Not started -- needs section 6's
- toggle first.
-- `match_play.py` gate clears `LOWER95 >= 0.5`. Not started, same
- dependency.
+ regression. **DONE** -- bit-identical solve counts on all three
+ (10/11, 83/90, 25/90), same-commit asm-vs-`GETATTACKS_BITBOARD=1` A/B
+ rather than vs. a now-stale `head_reference/` (see section 4 item 5
+ for why and the full numbers) -- also +8.38% aggregate wall-clock NPS.
+- `match_play.py` gate clears `LOWER95 >= 0.5`. Not started --
+ deliberately deferred for now, still required before retirement.
- `head_reference/` rebuilt as the new baseline once landed.
One easy win to check first, independent of all this: confirm whether