summaryrefslogtreecommitdiff
path: root/src/board_representation/EVAL.md
blob: 697ea380c4cb5e4930676d766e4326cf3a497fec (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
# Migration plan: bitboard-backed `Eval()` (rewrite, 2026-09-04)

**Status: planning only, superseding the earlier draft of this
document in full.** That draft was written before the movegen
bitboard work (`MOVEGEN_MIGRATION.md`) landed and assumed a slider
attack-bitboard primitive ("`_RookAttacksBB(c, bbOccupied)`... already
built") that did not actually exist yet at the time. It now does --
this rewrite is based on having read `_EvalBishop`/`_EvalKnight`/
`_EvalRook`/`_EvalQueen` in full (not by memory/analogy) and the
landed `generate.c` bitboard infrastructure (real magic tables, not a
sketch) side by side.

## 0. Why this rewrite exists, and what's different now

Two things changed since the first draft:

1. **Real magic bitboards now exist and are directly reusable.**
   `_RookAttacksBB(COOR c, BITBOARD bbOccupied)` / `_BishopAttacksBB(...)`
   (`generate.c`) are `FORCEINLINE`, non-static, and already used from
   three call sites outside their own file's move generation (the
   Part B check-block-mask computation, and would be a fourth caller
   here). They return the *complete* blocked-ray attack set in one
   multiply+shift+double-indirect-load -- not just the nearest
   blocker (`_WhoAttacksSquareBB`'s shape, built for a different
   question). This is exactly the primitive section 3a of the old
   draft assumed; it is no longer aspirational.
2. **The motivating problem is now measured, not assumed.** Profiling
   against Crafty on the same hardware (1.5Mnps vs. Crafty's 7.5Mnps,
   both at a similar lazy-eval skip rate) points at `Eval()` itself,
   not search, as the disproportionate cost. This plan exists to
   attack that number directly, not as a speculative "bitboards are
   modern, let's use them" exercise.

**Resolved 2026-09-05 (was an open question): yes, move `bbOccupied`
onto `POSITION`, incrementally maintained.** Investigated by reading
`move.c` end to end for how `bbPieces[2][8]`/`bbPawns[2]` -- the
precedent for exactly this kind of incrementally-maintained bitboard
-- actually get updated:

- Every mutation funnels through a small, fixed set of primitives
  (`SlidePiece`, `SlidePawn`, `LiftPiece`, `PlacePiece`, and their
  `WithoutSigs` variants) -- `MakeMove`/`UnmakeMove` never touch these
  bitboards directly. `bbOccupied` would touch the identical choke
  points, not a new set of call sites.
- It's actually *simpler* to maintain than `bbPieces`/`bbPawns`:
  those are keyed by piece type and color (an array lookup before the
  OR/AND), but occupancy doesn't care what's on a square or whose it
  is -- one unconditional clear-from-bit/set-to-bit pair per primitive,
  no branch, and it uniformly covers kings too (which today have no
  bitboard at all, only the `cNonPawns[.][0]` mailbox slot).
- **Currently it lives on `MOVE_STACK`** (`chess.h`'s `bbOccupied`
  field), rebuilt on demand via `_BuildFullOccupiedBB(pos)`
  (`generate.c`) -- itself already cheap (11 ORs of `bbPieces`/
  `bbPawns`/king squares, no mailbox scan, per that function's own
  comment) -- but called fresh at many independent sites: several in
  `generate.c`, four in `movesup.c`, one in `see.c`. Worse, the exact
  same 11-OR logic is duplicated verbatim as `_BuildOccupiedBB`
  (static, `see.c`) and `_BuildFullOccupiedBB` (non-static,
  `generate.c`) -- same function, two names, two files.
- Moving it to `POSITION`, incrementally maintained, means every one
  of those call sites (plus this plan's own future `eval.c` mobility
  rewrite, which needs the same value for
  `_RookAttacksBB`/`_BishopAttacksBB(c, bbOccupied)`) reads one
  already-current field instead of independently re-deriving it --
  and dedupes the two identical builder functions into one.
- **Verification is close to free**: `generate.c` already has
  `ASSERT(pStack->bbOccupied == _BuildFullOccupiedBB(pos))` at three
  call sites, cross-checking the `MOVE_STACK`-cached copy against a
  from-scratch rebuild today. The identical assert, repointed at
  `pos->bbOccupied`, becomes the DEBUG-build safety net for the
  incremental version -- the same pattern that already validated
  `bbPieces`/`bbPawns` when they were added.
- **Scoped as standalone, low-risk work, not gated on the rest of
  this plan**: move generation benefits from it immediately regardless
  of `Eval()`'s progress, so it's worth landing (add the field, touch
  the handful of `move.c` primitives, dedupe the two builder
  functions, repoint the existing assert, run
  `debug_smoke_test.sh`/`precommit_check.sh`) on its own, ahead of or
  alongside section 4's toggle work -- not bundled into any single
  piece type's toggle.

**The movegen project's own findings are the load-bearing precedent
here, and they cut both ways -- worth stating plainly before
proposing more work of the same shape:**

- Most individual `_Generate*BB` functions landed at **speed parity**
  with mailbox (0.71x-1.10x depending on position density), because a
  mailbox ray walk's per-square cost was already close to
  O(destination count) -- nothing wasted to reclaim.
- The **real, consistent wins** were in a different place: the
  dispatch layer (`_GenerateAllMovesBB` avoiding `JumpTable`'s
  indirect, badly-predicted call, up to 23% in dense positions) and
  `IsAttackedBB` (0.73x-0.93x, a genuine algorithmic improvement, not
  just a constant factor, because mailbox `IsAttacked` is O(piece
  count) and the bitboard version is closer to O(1)).
- The stashed *first* Eval bitboard attempt regressed 17-41% for a
  root cause fully diagnosed at the time: **pure duplication** --
  every mobility function ran its full original ray-walk *and*
  additionally wrote bitboards, so every position paid for both
  representations and nothing was ever removed to pay for it. Not
  evidence against this approach; evidence against ever doing it
  additively again.

Reading `Eval()`'s mobility loops against this precedent suggests
`Eval()` is actually a *better* candidate for a bitboard win than move
generation was, for a structural reason move generation never had:
**every mobility loop today pays for an indirect `switch` dispatch per
square visited** (`RMobCaseTable`/`BMobCaseTable`/`QMobCaseTable`/
`NMobCaseTable`), explicitly chosen over a function-pointer jump table
"to avoid an indirect call/ret... since the target piece varies square
to square" (`_EvalBishop`'s own comment). A `switch` compiles to a
jump table or branch chain that itself suffers exactly the
misprediction problem `_GenerateAllMovesBB` was built to eliminate
for move generation's dispatch layer -- and it runs *inside* the
hottest loop in eval, once per square walked, not once per node. This
plan's central bet is that removing this dispatch, not the ray-walk
arithmetic itself, is where the win lives -- directly analogous to
what actually paid off in the movegen project, not the part that
didn't.

## 0b. Performance philosophy for this rewrite -- what's negotiable
and what isn't

Stated by the user directly, 2026-09-05, and important enough to
record verbatim as a standing constraint rather than let it live only
in chat history: this engine runs at ~1.5Mnps against Crafty's
~7.5Mnps on the same hardware, both profiled, and `Eval()` is the
confirmed disproportionate cost -- not search. That changes the
default posture for every decision in this plan:

- **Benchmark everything, cut aggressively.** Every existing eval
  term is a candidate for removal if it doesn't earn its cost --
  section 7's cut list (dead `#if 0` code, max-mobility-in-a-row
  bonus, bishop's transient-pawn credit, connected-rook x-ray
  bookkeeping) should be treated as the *starting* set of suspects,
  not an exhaustive one, and the DNA-zeroing sensitivity check +
  cycle-cost measurement pairing should be run proactively across
  `Eval()`'s terms rather than opportunistically on ones that already
  look suspicious.
- **The one thing that is not up for negotiation: attack-bitmap-driven,
  mobility-and-safety-aware piece evaluation itself** -- not merely
  "king safety," a narrower framing worth explicitly correcting here.
  What today's `bvAttacks`/case-table machinery gives every piece
  type is a single, unified mechanism for "does this piece actually
  have useful, safe mobility given the current pawn structure and
  attack picture" -- bad bishops, knights needing real outposts (not
  just central *squares*, since a centralized knight throttled by an
  enemy pawn chain gets no credit), rooks on genuinely open/contested
  files, and king danger, all falling out of the same attack-bitmap
  substrate rather than being separate hand-tuned heuristics. In the
  user's own words: *"If I had to drop everything else from Eval to
  afford this, I would still keep it. I believe it is what chess Eval
  is about."* Section 2's replacement (`bbPawnAttacks`/
  `bbMinorAttacks`/`bbRookAttacks`/`bbQueenAttacks` accumulators) is
  in scope and *encouraged* precisely because it's the fast
  replacement for this capability's current slow mechanism -- the
  capability survives, only its implementation gets cheaper. What
  must not happen is a cut that removes mobility/safety-awareness
  itself in the name of speed (e.g. reverting a piece type to
  raw material + PST with no attack-picture-conditioned mobility term
  at all) -- every other term in `Eval()` is fair game for that kind
  of cut; this one specifically is not.
- **Profile before rewriting, not just before/after each toggle.**
  Before continuing further piece-type work, get a per-eval-term
  cycle-cost breakdown (perf counters or sampling, representative
  position mix) of where `Eval()`'s total time actually goes today --
  mobility ray-walks vs. pawn structure vs. king safety vs. everything
  else. This should drive both rewrite sequencing and cut-candidacy;
  section 4's knight-first ordering was sequenced by implementation
  risk (fewest special cases), which is a reasonable tie-breaker but
  shouldn't override what the actual profile says is worth attacking
  first once that data exists.

## 0c. `_EvalPawns` -- measured cost, and a concrete redundant-work
finding (2026-09-05)

The `EVAL_TIME` per-term instrumentation added per section 0b (see
`chess.h`'s `u64CyclesEval*` counters, `root.c`'s breakdown print)
gave a first real number: on one representative middlegame position
(`sd 12`), `_EvalPawns` alone was **5.4% of total cycles spent in
`Eval()`**, despite a measured pawn-hash hit rate above 99% -- i.e.
this cost is overwhelmingly the *hit* path (hash probe + key compare
+ return), not the rare rebuild-on-miss path. Worth remembering for
section 0b's "front-load if cheap, else estimate" framing: pawn
structure is the *first* thing evaluated in `Eval()` today, which is
right for a term this correlated with lazy-eval accuracy, but "first"
doesn't mean "free" -- even a >99%-hit-rate cache probe measurably
adds up at billions of calls, same lesson as the mobility dispatch
switch in section 0.

**Concrete, scoped finding on the miss path itself** (found by reading
`_EvalPawns`, eval.c:1866 on, against `POSITION`'s own fields): on a
pawn-hash miss, `_EvalPawns` rebuilds `pHash->bbPawnLocations[uColor]`
bit-by-bit inside its per-pawn loop (`pHash->bbPawnLocations[uColor]
|= COOR_TO_BB(c);`, eval.c:1958) by iterating `pos->cPawns[uColor][u]`
-- but `POSITION` already carries `pos->bbPawns[2]`
(`chess.h`:661-670), a plain, always-current, incrementally-maintained
pawn-location bitboard (`move.c` updates it on every pawn move,
`fen.c`/`board.c` build it at position-load time), entirely separate
from the pawn-hash-keyed `pHash->bbPawnLocations[2]`. There's no
reason for the miss path to reconstruct from scratch what's already
sitting on `pos`: `pHash->bbPawnLocations[uColor] =
pos->bbPawns[uColor];` once per color replaces the bit-by-bit OR
inside the loop. **The loop itself still has to run** for the
per-file counts (`pHash->uCountPerFile`) and whatever isolated/
doubled/duo/passed-pawn detection follows -- this only removes the
bitboard-population part of that loop's work, not the loop -- and
since it's gated behind the <1%-of-calls miss path, the aggregate win
is real but necessarily small. Flagged here as a correctness-safe,
low-risk cut to take regardless of the rest of this plan's sequencing,
not because it's expected to move the 5.4% number much on its own.

**Explicit scope note, stated by the user directly**: `_EvalPawns`
stays pawn-structure-only -- passed-pawn detection, connectivity, and
whatever else it currently does for pawns are all still wanted. The
ask here is narrower than "simplify pawn eval": stop redoing bitboard
work `POSITION` already maintains for you, don't cut pawn eval terms.

Every one of `_EvalBishop`/`_EvalKnight`/`_EvalRook`/`_EvalQueen` was
read end to end for this rewrite (`eval.c`). Two clearly separable
halves exist in every one of them:

### 1a. Non-mobility terms -- already fine, not a target

Open/half-open file bonuses, rook-behind/leads-passer, knight
outpost/tropism, bishop good/bad-pawn-color counting, centrality,
closed-position scalers -- all of these read `pHash`-cached bitboards
(`bbPasserLocations`, `bbStationaryPawns`, `bbPawnLocations`,
`uCountPerFile`) or O(1) `POSITION` fields. No ray-walk, no per-square
switch, already bitboard-driven where it needs to be. **Do not touch
these** -- rewriting them buys nothing and only adds risk surface.

### 1b. Mobility ray-walks -- the actual target, and the switch is
the point, not the walk

All four piece types share the same shape: walk each ray direction
square by square, at every square (1) OR a bit into
`pos->rgSquare[cSquare|8].bvAttacks[uColor]` (see section 2), (2)
dispatch on `pos->rgSquare[cSquare].pPiece` via a per-color 14-entry
case table to decide "count this square? keep walking? x-ray past
it?". Read closely, the case tables collapse to a small, enumerable
set of primitive facts once you stop thinking per-square and start
thinking per-ray-segment:

**Rook (`RMobCaseTable`, 6 live cases):**
- `RMOB_EMPTY` / `RMOB_ENEMY_LESS` (any enemy piece, actually --
  "less" is a stale name, the table just needs "count if safe, stop
  if occupied"): count if `!UNSAFE_FOR_ROOK`, stop if occupied.
- `RMOB_FRIEND_BLOCK` (friendly non-rook/queen): stop, no count.
- `RMOB_FRIEND_ROOK` / `RMOB_FRIEND_QUEEN`: don't stop -- x-ray past a
  same-color battery partner, no mobility credit at the blocker
  itself, award `ROOK_CONNECTED_HORIZ`/`_VERT` at the blocker square.
- `RMOB_ENEMY_SAME` / `RMOB_ENEMY_GREATER`: count (unconditionally,
  no unsafe-check -- this is deliberate, capturing an equal/higher
  rook/queen/king is never "unsafe" in the sense mobility cares
  about), stop.

Reduction: `bbAttack = _RookAttacksBB(c, bbOccupied)` already stops at
the nearest blocker in every direction by construction -- that
blocker is exactly the one square each case above is examining.
`uTotalMobility = CountBits((bbAttack & ~bbFriendlyOccupied &
~bbUnsafeForRook) | bbEnemyBlockerCredit)`, where `bbEnemyBlockerCredit`
handles the fact that an occupied-by-enemy terminal square counts
*without* the unsafe check (`RMOB_ENEMY_SAME`/`_GREATER` never test
`UNSAFE_FOR_ROOK`, only `RMOB_EMPTY`/`_ENEMY_LESS` -- i.e. the
non-terminal-and-terminal-enemy cases both existed in the original
table, but only the *empty-square* case is masked by
`bbUnsafeForRook`; every enemy-occupied terminal square counts
unconditionally regardless of case name). Concretely:

```c
bbAttack    = _RookAttacksBB(c, bbOccupied);
bbEnemy     = bbAttack & bbEnemyOccupied;      /* terminal enemy squares */
bbEmpty     = bbAttack & ~bbOccupied;          /* non-terminal empty squares */
bbMobility  = bbEnemy | (bbEmpty & ~bbUnsafeForRook);
uTotalMobility = CountBits(bbMobility);
```

The friendly-rook/queen x-ray case needs its own handling since a
magic lookup's attack set stops *at* the blocker regardless of type,
giving no visibility past it: detect `bbFriendRQ = bbAttack &
(pos->bbPieces[uColor][ROOK] | pos->bbPieces[uColor][QUEEN])`
(0 or 1 bits, essentially always), and only when nonzero, recompute
`_RookAttacksBB(c, bbOccupied & ~bbFriendRQ)` restricted to the same
ray (`g_RookRayToEdge[dir][c]`) to get the far-side squares for the
attack-bit population in section 2 -- this does not change
`uTotalMobility` (the original code never credited mobility past a
friendly battery partner either). The connected-rook bonus itself
(`ROOK_CONNECTED_HORIZ`/`_VERT`) needs to know if the ray to the
blocker was horizontal or vertical, which is already known for free
from *which* of the 4 ray directions found it.

**Per-direction max mobility**: `CountBits(bbMobility &
g_RookRayToEdge[dir][c])` for each of 4 directions -- 4 extra popcounts,
still branch-free, no walk.

**Bishop (`BMobCaseTable`, 7 live cases) -- one genuine wrinkle rook
doesn't have:**

- `BMOB_EMPTY` / `BMOB_ENEMY_PAWN`: same shape as rook's
  empty/enemy-less, masked by `UNSAFE_FOR_MINOR`.
- `BMOB_FRIEND_BLOCK`: stop, no count (friendly knight/rook/king).
- `BMOB_FRIEND_PAWN`: stop, **but still counts as 1 mobility square
  if `pos->bb & COOR_TO_BB(cSquare)`** -- `pos->bb` here is a
  *scratch* alias set to "pawns of either color, non-stationary,
  restricted to this bishop's own color complex" just before the walk
  begins (`pos->bb = bbPc`, see the code right above the ray loop).
  This is real, intentional signal (a friendly *transient* pawn -- not
  rammed/backward -- sitting on the bishop's own diagonal still
  "counts" toward mobility, on the theory it'll likely move and open
  the diagonal soon) and is **the one case that isn't a pure
  `~bbFriendlyOccupied` mask** -- it needs `bbAttack`'s terminal
  friendly-pawn bit added back in when that pawn is transient and on
  the bishop's own color. Since at most one square per ray direction
  can be this case (the walk stops there), this is a cheap
  `bbAttack & pos->bbPawns[uColor] & bbTransientOwnColorPawns` term
  ORed into the mobility mask, not a structural problem -- just a
  named exception that must not be dropped silently during the
  rewrite.
- `BMOB_ENEMY_SAME` (opposing bishop/knight): count, stop.
- `BMOB_FRIEND_XRAY` (friendly bishop/queen): x-ray past, no count,
  same shape as rook's battery case.
- `BMOB_ENEMY_GREATER` (enemy rook/queen/king): **count, then x-ray
  past anyway** -- unlike rook, bishop's x-ray set includes both the
  same-type friendly case *and* an enemy-major-piece case. This means
  the "recompute with blocker excluded" trick (needed for attack-bit
  population past the blocker, section 2) has to run for *two*
  distinct terminal-piece categories here, not one.

**Queen (`QMobCaseTable`, 7 live cases) -- mechanically rook+bishop,
with one extra wrinkle from combining two ray families:**

- Same empty/enemy-less/friend-block/enemy-same shape as rook, masked
  by `UNSAFE_FOR_QUEEN`.
- `QMOB_FRIEND_BISHOP` x-rays only on a diagonal ray, blocks on an
  orthogonal one (`fStop = fOrthogonalRay`); `QMOB_FRIEND_ROOK` is the
  mirror (`fStop = !fOrthogonalRay`). This is already exactly what
  computing queen mobility as two separate rook-direction/
  bishop-direction magic lookups naturally gives for free -- no new
  logic needed, `fOrthogonalRay`'s role is entirely subsumed by "which
  of the two lookups this ray direction belongs to."
- `QMOB_FRIEND_QUEEN`: x-ray past (either ray family), same as rook's
  own-type case.
- `QMOB_ENEMY_GE` (any enemy piece): count, stop, **no x-ray past** --
  unlike bishop's `BMOB_ENEMY_GREATER`, queen does not x-ray through a
  captured-but-higher-value enemy. One fewer special case than bishop,
  not more.

Reduction: `bbAttack = _RookAttacksBB(c, occ) | _BishopAttacksBB(c,
occ)`, each masked/unsafe-checked with its own family's mask
(`bbUnsafeForQueen` uniformly, since the mailbox code doesn't
distinguish rook-direction vs. bishop-direction safety for the
queen), matching `MOVEGEN_MIGRATION.md` section 3's own already-tested
"combined 8-ray table measured slower than reusing the rook/bishop
tables in two passes" finding -- reuse that two-pass structure here
too, don't rediscover the same regression.

**Knight (`NMobCaseTable`, 3 live cases) -- the simple case, confirmed
by direct read:**

- `NMOB_MOBILE_SQUARE` (empty or enemy pawn): count if
  `!UNSAFE_FOR_MINOR`.
- `NMOB_ENEMY_OTHER` (any other enemy piece): count, unconditionally.
- `NMOB_FRIEND`: no count.

No terminal-blocker subtlety at all (knights don't block through
anything). Reduction is exactly one line:

```c
bbAttack   = g_KnightAttacksBB[c];
bbMobility = (bbAttack & bbEnemyNonPawnOccupied) |
             ((bbAttack & ~bbFriendlyOccupied & ~bbEnemyNonPawnOccupied)
              & ~bbUnsafeForMinor);
uMobilitySquares = CountBits(bbMobility);
```

(the split exists only because `NMOB_ENEMY_OTHER` skips the unsafe
check and `NMOB_MOBILE_SQUARE` doesn't -- can likely simplify further
once `bbEnemyNonPawnOccupied`'s exact membership is nailed down
against the case table above).

### 1c. What every one of these functions does *after* mobility --
unaffected

Trapped-piece recording (`_RecordTrappedCandidate`, gated on
`uTotalMobility == 0`/`< 3`) and the rook-corner king-trap check are
already independent of *how* the mobility count was computed --
feed them the bitboard-computed number, no rewrite needed.

## 2. `bvAttacks` replacement -- still the load-bearing new
infrastructure, design unchanged from the first draft, now
concretely groundable against real code

`ATTACK_BITV` (`chess.h`) packs presence bits per piece-type family
into one `ULONG` per square per color half
(`pos->rgSquare[c|8].bvAttacks[color]`), with `UNSAFE_FOR_MINOR`/
`_ROOK`/`_QUEEN` macros masking the relevant bits
(`0x80`/`0xC0`/`0xE0` -- i.e. each wider piece type's "am I safe here"
question is a superset of the narrower one's, matching chess reality:
a queen cares about more attackers than a rook, which cares about more
than a minor). This is read constantly (every mobility loop's
`UNSAFE_FOR_X` check) and written constantly (every mobility loop's
`|= uBit`), in a strict piece-evaluation order that later pieces
depend on.

Replacement, same shape as the first draft proposed, now stated
precisely against the actual bit semantics above:

- `Eval()`-local (not `POSITION`-resident -- same call-scoped-lifetime
  discipline `MOVEGEN_MIGRATION.md` section 3c/6b already established
  for its own scratch bitboards) per-side accumulators:
  `bbPawnAttacks[2]`, `bbMinorAttacks[2]` (knight | bishop, matching
  `UNSAFE_FOR_MINOR`'s own single-tier mask), `bbRookAttacks[2]`,
  `bbQueenAttacks[2]`.
- Populated in `Eval()`'s existing piece-evaluation order (pawns,
  then knights/bishops, then rooks, then queens, then king), each
  piece type OR-ing its own `bbAttack` (the same value computed for
  its own mobility, no second computation) into its side's
  accumulator once, immediately after that piece's own mobility is
  scored.
- `bbUnsafeForMinor` at knight/bishop-evaluation time =
  `bbPawnAttacks[enemy]` (matches the `0x80` mask -- only pawns make a
  square unsafe for a minor). `bbUnsafeForRook` = `bbPawnAttacks[enemy]
  | bbMinorAttacks[enemy]` (`0xC0`). `bbUnsafeForQueen` adds
  `bbRookAttacks[enemy]` (`0xE0`). Each is a single OR of
  already-populated accumulators, computed once per side per `Eval()`
  call -- not once per candidate square, a strict improvement over
  today's per-square `ATTACK_BITV` read.
- **King safety zone queries** (`_EvalKing`'s read side, section 3
  below) become `CountBits(bbXAttacks[enemy] & kingZoneMask)` directly
  against these same accumulators -- no separate data structure
  needed for that consumer.

One correction versus the first draft's framing: `pos->rgSquare[]`
itself is **not fully retirable** even after this lands, for a reason
the movegen project's own `movesup.c` survey already surfaced --
every generator, bitboard or not, still needs `pos->rgSquare[c].pPiece`
to answer "what's on the terminal/blocker square" (for the
enemy-vs-friendly, same-vs-greater-value classification every piece
type's reduction above still needs at the *one* terminal square per
ray, not for the whole ray). This plan retires the mailbox *ray-walk
and the bvAttacks per-square bitfield*, not the mailbox array itself.

## 3. `CountKingSafetyDefects`/`_EvalKing` -- unchanged scope from the
first draft, still explicitly separate

`_EvalKing`'s *read* side (walking `KingSafetyDeltas` around the
king, summing attacker/defender presence) becomes a `CountBits`
query against section 2's accumulators, a pure consumer change.
`CountKingSafetyDefects` itself is **not** a port target -- per the
first draft's already-recorded decision, it gets a from-scratch
bitboard-native reimplementation that need not preserve exact
semantics (allowed to be *more accurate*, since it's a lazy-eval
magnitude hint, not a scored term itself), scoped and gated entirely
separately from the mobility rewrite below. Do not bundle it into the
same toggle or the same benchmark pass.

**Correction, found by reading the actual call sites (not by
analogy): `CountKingSafetyDefects` cannot consume section 2's
accumulators even opportunistically, and its role is bigger than "a
lazy-eval magnitude hint."**

1. **Ordering makes reuse structurally impossible, not just
   undesirable.** `CountKingSafetyDefects` runs during `Eval()`'s
   lazy-eval margin phase (`EstimatePositionalScore`, and the
   `CountKingSafetyDefects(pos, WHITE/BLACK)` refresh calls at
   eval.c:4976-4977) -- strictly *before* the piece-eval loop that
   builds `bvAttacks` today, and before section 2's `bbPawnAttacks`/
   `bbMinorAttacks`/`bbRookAttacks`/`bbQueenAttacks` accumulators
   would exist under this plan (pawns -> knights/bishops -> rooks ->
   queens -> king). There is no rewrite of `CountKingSafetyDefects`
   that can read those accumulators; by construction they aren't
   populated yet at the point it's called. It has to remain what it
   already is: a geometry-only estimate over enemy piece *locations*
   relative to the king (`CHECK_VECTOR_WITH_INDEX`), with no attack
   data available, bitboard-native or otherwise.
2. **It is not only a lazy-eval input.** `search.c:1252`
   (`CountKingSafetyDefects(pos, pos->uToMove) > 2`) and
   `searchsup.c:484` (`CountKingSafetyDefects(&ctx->sPosition, uColor)
   > 1`) call it directly as a search-time extension/reduction gate,
   fully independent of `Eval()`'s lazy-exit path. A drift here
   doesn't just blur a margin estimate -- it changes which nodes get
   extended or reduced at every level of the tree.

**The actual constraint this plan needs, stated precisely**: not
"port `CountKingSafetyDefects` to bitboards," but *"whichever accuracy
improvements section 3's `_EvalKing` rewrite introduces (e.g. properly
catching x-ray/latent queen threats via the new accumulators),
`CountKingSafetyDefects`'s cheap geometric estimate must remain
correlated with `_EvalKing`'s real danger score across the position
space."* Concretely, the two thresholds already load-bearing in
production (`> 1` in `searchsup.c`, `> 2` in `search.c`) need to keep
firing on roughly the same set of positions `_EvalKing`'s new,
possibly-more-accurate score would flag as dangerous. If `_EvalKing`
gets meaningfully better while `CountKingSafetyDefects` stays exactly
as it is today, that correlation can silently degrade even though
neither function, read in isolation, has a bug -- the failure mode is
between them, not inside either one.

**Added verification step (section 5 should include this, not just
`_EvalKing`'s own exact-score harness)**: over the same 20,000-position
`GenerateRandomLegalPosition` sample, compute both
`CountKingSafetyDefects(side)` (old, unchanged) and the new
`_EvalKing`'s real danger score, old vs. new, and check:

- **Correlation, not equality** -- these were never meant to match
  exactly (one is a cheap O(pieces) geometric proxy, the other a full
  bitboard-weighted computation); the thing to track is whether their
  relative ranking of "how dangerous is this king position" stays
  consistent before vs. after the `_EvalKing` rewrite.
- **Threshold agreement specifically at the two live gates** -- for
  each sampled position, does `CountKingSafetyDefects(side) > 1` (and
  `> 2`) still agree with "old `_EvalKing` considered this position
  dangerous" at roughly the same rate it agrees with "new `_EvalKing`
  considers this position dangerous"? A meaningful shift in agreement
  rate, not just a raw score delta, is the signal that
  `CountKingSafetyDefects` itself may need re-tuning (its
  `KingFlightDefects`/`KingStormingPawnDefects`/`CHECK_VECTOR`
  constants are DNA-tunable, per `eval_tune/`) to stay a useful proxy
  for the new, more accurate ground truth -- not evidence that the
  `_EvalKing` rewrite itself is wrong.
- **Downstream check**: re-run `match_play.py` and the curated-suite
  `sd10` gate (section 5.2/5.3) with `_EvalKing`'s rewrite landed but
  *before* touching `CountKingSafetyDefects`'s constants, specifically
  watching for search-behavior regressions (unexpected extension/
  reduction pattern shifts) that a pure eval-score comparison would
  miss, since sections 5.1-5.3 as currently scoped only compare
  `Eval()`'s output, not `CountKingSafetyDefects`'s search-time
  pruning effect.

## 4. Toggle strategy -- learn from the stash's mistake explicitly

One `#define` per piece type (`EVAL_KNIGHT_BITBOARD`/
`EVAL_BISHOP_BITBOARD`/`EVAL_ROOK_BITBOARD`/`EVAL_QUEEN_BITBOARD`),
matching the movegen project's per-piece-type convention -- **but each
toggle must replace that piece type's entire mobility+attack-bit
computation, never add a bitboard computation alongside the mailbox
one.** This is not a style preference; it is the single, fully
diagnosed cause of the stashed attempt's 17-41% regression, and must
be treated as a hard constraint on every toggle in this plan, not a
lesson to merely keep in mind. Section 2's `bbUnsafeFor*`
infrastructure is shared prerequisite state, needed before *any*
piece type's toggle can be flipped on -- there's no meaningful
partial-rollout ordering where accumulator infrastructure lands after
a consumer.

Suggested order, cheapest/most-precedented first (matches
`MOVEGEN_MIGRATION.md`'s own successful sequencing rationale):
knight (no terminal-blocker special case at all) -> rook (one x-ray
category) -> bishop (two x-ray categories, one non-mask-reducible
transient-pawn credit) -> queen (mechanical once rook+bishop are
proven, reuses their two-pass structure).

## 5. Correctness verification

Same bar as `MOVEGEN_MIGRATION.md`, arguably higher since `Eval()`
feeds every leaf score in the tree rather than just node
ordering/generation:

1. **Exact-score comparison harness**: for each of
   `GenerateRandomLegalPosition`'s 20,000-position sample (same source
   the movegen project's own harness used), compute the specific
   `EVAL_TERM` contributions this plan touches (mobility +
   connected/x-ray bonuses + trapped-candidate flag) old vs. new,
   per piece type independently (mirrors the movegen project's
   per-piece-type toggle granularity in the harness too). Must be
   byte-identical -- `Eval()` is deterministic-by-design, any mismatch
   is a real bug.
2. **Whole-engine `sd10` on all three curated suites vs.
   `head_reference`** -- rebuilt post-movegen-landing baseline, so
   this plan's own gate isn't contaminated by the movegen commit's
   own (already-confirmed-unrelated) `ecm_hard_quick` drift.
3. **`match_play.py`** (`LOWER95 >= 0.5`), same gate, same reasoning:
   given the movegen project's own experience that most individual
   functions land near parity, expect this gate to converge toward
   "no regression" rather than a large positive swing, and don't
   over-read a near-0.5 LLR as a problem the way the movegen SPRT run
   didn't need to either.
4. **`precommit_check.sh`** as always, for the crash/assert layer.

## 6. Microbenchmarking

Same two-tier approach as every prior plan in this family:

1. **Isolated cycles/call per piece type**, interleaved old/new,
   across opening/middlegame/endgame density, modeled on
   `testgenerate.c`'s existing methodology (new `testeval.c`-resident
   harness, or extend the existing one). Given section 0's
   dispatch-elimination bet, pay particular attention to whether the
   win (if any) tracks with *how many pieces/squares* a position has
   to walk (more squares -> more switch mispredictions saved) rather
   than being flat -- a density-correlated win would confirm the
   dispatch-removal theory; a flat or absent one would suggest the
   `switch` was actually being predicted fine and the bet was wrong.
2. **Whole-`Eval()` cycles/call**, same interleaved technique -- the
   number that actually matters, since mobility is only part of
   `Eval()`'s total cost.
3. **Whole-engine NPS** on the curated suites, and directly against
   the original motivating comparison: **re-run the same Crafty
   side-by-side NPS comparison that motivated this plan**, not just
   an internal before/after -- the actual goal is closing (some of)
   the 1.5M-vs-7.5Mnps gap, and that number is the one that should
   ultimately justify the effort here, not an isolated microbenchmark
   in a vacuum.

## 7. What to consider cutting, not just speeding up

Raised at the user's request -- anything that resists a clean bitboard
reduction is a candidate for this discussion, not an automatic
"implement it the hard way":

- **Dead code, delete regardless of this plan**: every `#if 0`
  min-mobility-tracking block in `_EvalRook`/`_EvalBishop`/
  `_EvalQueen`/`_EvalKnight` -- unused (`pos->uMinMobility` is written,
  never read live), pure removal, zero risk, zero dependency on
  anything else in this plan.
- **Per-direction max-mobility bonus**
  (`ROOK_MAX_MOBILITY_IN_A_ROW_BONUS`/`BISHOP_MAX_MOBILITY_IN_A_ROW_BONUS`) --
  survives the rewrite as 4 extra `CountBits` calls per slider, cheap
  either way, but worth a DNA-tuning-style sensitivity check (does
  zeroing this table's contribution move solve rate or match score at
  all?) independent of the bitboard work -- if it's not pulling
  weight, cutting it removes both eval-time cost and tuning-parameter
  count for free. Not blocking this plan; worth doing opportunistically
  once the rewrite is in place and the harness exists to check it
  cheaply.
- **Bishop's `BMOB_FRIEND_PAWN` transient-pawn mobility credit**
  (section 1b) -- the single case that doesn't reduce to a mask
  operation as cleanly as everything else. It's cheap to keep (one
  extra AND term), but if the isolated-benchmark step shows bishop's
  reduction costing more than knight/rook/queen's for a
  disproportionately small eval-quality contribution, this specific
  case is the one to question first, precisely because it's the one
  genuinely bespoke piece of logic in the whole rewrite. Measure
  before deciding, per section 6 -- don't cut preemptively.
- **`CountKingSafetyDefects`** -- not a "cut" candidate, but flagged
  again here because it's the one function in this whole area that
  isn't a straightforward port either way (section 3); if the
  from-scratch reimplementation turns out to be disproportionately
  expensive to get right relative to its actual scoring impact, that's
  a legitimate place to ask "is a simpler, less-accurate hint good
  enough" rather than chasing full replacement fidelity.
- **Connected-rook / bishop-x-ray attack-bit population past a
  blocker** (section 1b's "recompute occupancy with blocker excluded"
  trick) -- this exists purely to keep populating `bvAttacks`-derived
  king-safety-relevant bits correctly for squares beyond a battery
  partner. If, once section 2's accumulators are live and
  `_EvalKing`'s consumption is measured, this far-side attack
  information turns out to contribute negligibly to king-safety
  scoring in practice (worth checking via the same DNA-sensitivity
  technique), it's a legitimate cut -- dropping it would remove the
  single most annoying piece of bookkeeping in this entire plan (the
  two-terminal-category case for bishop, section 1b) for what might be
  a very small eval-quality cost. Explicitly flagged as the highest-
  complexity-per-value item in the plan; measure before committing to
  full fidelity.

## 8. Retirement criteria

Same shape as `MOVEGEN_MIGRATION.md` section 7: per piece type, delete
that type's mailbox mobility ray-walk only after the exact-score
harness (section 5.1), the `sd10` curated-suite check (section 5.2),
the isolated-cycles benchmark (section 6.1) showing a consistent win
(or an explicit decision to keep it at parity for the dispatch-layer
win, same reasoning the movegen project used for its own near-parity
functions), and `match_play.py` (section 5.3) all pass. Don't let
knight's likely-clean bill of health lower the bar for bishop's
two-x-ray-category case or queen's combined-ray-family case --
each has different enough special-case surface (section 1b) to
warrant its own full pass through this list.