After 1b, triangulating both polygons on every pair was the largest
remaining overlap cost (27% of main-thread samples on the corpus job).
PartOverlapChecker now triangulates each part at most once per check,
lazily after the bounding-box gate, and passes the triangles to a new
internal Collision.HasOverlap overload that runs the unchanged
OverlapRegions body. Triangles are only read by clipping and hole
subtraction, so reuse gives identical verdicts.
Verification:
- 49,000 seeded decisions with reused triangles match LegacyCollision;
triangles stay bit-identical to a fresh triangulation afterwards.
- Debug PolygonTriangulations: 246 -> 40 and 64 -> 36 per grid check;
sharing triangles per Program instead fails 23 tests.
- Corpus job (169 parts, --engines Default --parallel 1): median
13,398 -> 12,702 ms over 4+4 alternating runs vs 1b, identical
outcomes; serialized layout byte-identical to the base.
Also records the Follow-up B' (Slices 1a, 1b, 2a) measurements in
docs/performance/fill-performance.md.