Commit Graph
8 Commits
Author SHA1 Message Date
aj 1b23ad79f2 perf(fill): reuse part triangulations within an overlap check
After 1b, triangulating both polygons on every pair was the largest
remaining overlap cost (27% of main-thread samples on the corpus job).
PartOverlapChecker now triangulates each part at most once per check,
lazily after the bounding-box gate, and passes the triangles to a new
internal Collision.HasOverlap overload that runs the unchanged
OverlapRegions body. Triangles are only read by clipping and hole
subtraction, so reuse gives identical verdicts.

Verification:
- 49,000 seeded decisions with reused triangles match LegacyCollision;
  triangles stay bit-identical to a fresh triangulation afterwards.
- Debug PolygonTriangulations: 246 -> 40 and 64 -> 36 per grid check;
  sharing triangles per Program instead fails 23 tests.
- Corpus job (169 parts, --engines Default --parallel 1): median
  13,398 -> 12,702 ms over 4+4 alternating runs vs 1b, identical
  outcomes; serialized layout byte-identical to the base.

Also records the Follow-up B' (Slices 1a, 1b, 2a) measurements in
docs/performance/fill-performance.md.
2026-09-27 13:49:46 -04:00
aj a27290a29c perf(fill): prepare overlap polygons once per check
Both HasOverlappingParts loops rebuilt each part's polygon from its
Program on every pair. PartOverlapChecker prepares each distinct Program
(reference identity) once and each part's world polygon once per call,
then uses the overlap-only Collision.HasOverlap. Loop order, bounding-box
prefilter, early exit and returned indices are unchanged; Part.Intersects
shares the material/polygon recipe and still returns crossing points.

Verification:
- Frozen LegacyPartOverlap differential (original Intersects and both
  loops): verdicts, indices and world polygons bit-identical across fill
  grids, patterns, touching/epsilon gaps, scribe/rapid/empty programs.
- Debug OverlapPolygonPreparations: 246 -> 1 and 64 -> 2 per check.
- Corpus job (169 parts, --engines Default --parallel 1, with 1a):
  median 18,885 -> 13,464 ms over 4+4 alternating runs, identical
  outcomes; serialized layout byte-identical to the base.
2026-09-27 13:34:04 -04:00
aj 2b5485f6cf perf(core): skip unused crossing points in overlap-only checks
Collision.HasOverlap only needs the verdict, but it went through Check,
which also collected crossing points. Triangulation, clipping and hole
subtraction now live in one private OverlapRegions method shared by Check
and HasOverlap, so verdict arithmetic stays single-sourced; Check output is
unchanged.

Tests: a frozen copy of the previous Collision is the oracle. 50,000 seeded
HasOverlap verdicts and 2,400 bitwise Check results match it, plus
containment, contact, hole and input-immutability cases. A Debug-only
PerfCounters.CrossingPointScans counter proves HasOverlap no longer scans.
Malformed polygons with null outer vertices still throw when the bounding
boxes overlap (now ArgumentNullException from triangulation rather than
NullReferenceException from ToLines).

Measured (Release, same harness in both trees): about 44% less time per
overlap-only polygon check, allocations 10.0 -> 7.9 MB per 155-pair sweep.
The 169-part serialized corpus layout is byte-identical.
2026-09-27 12:54:20 -04:00
aj 8188533d72 perf(ml): support scalar-only angle features 2026-09-26 00:02:28 -04:00
aj 4553f8afad perf(fill): avoid redundant bounds recomputation 2026-09-25 21:45:03 -04:00
aj 6efa6b1117 perf(fill): remove discarded extents pitch geometry 2026-09-25 19:52:19 -04:00
aj b318950a54 perf(fill): short-circuit default comparisons by count 2026-09-25 16:43:09 -04:00
ajandClaude Opus 5.5 1c363504f5 perf(engine): key fill caches by source drawing so they hit across trials
Every DefaultPlateFiller.Fill makes a fresh canonical copy of the drawing,
and BestFitCache/FillResultCache keyed by drawing reference, so fills never
shared results and the static caches grew without bound.

CanonicalFrame now records which drawing each canonical copy came from.
Both caches key weakly on that source drawing, so every canonical copy
shares one entry and released drawings can be collected. An entry is
dropped when the drawing's Program instance or canonical angle changes.
Best-fit candidates are computed once per (drawing, spacing) through
BestFitFinder.FindCandidates and filtered per plate size with the same
filter FindBestFits uses. FillResultCache keeps canonical and
non-canonical callers apart.

Adds Debug-only PerfCounters for best-fit runs, offset perimeter builds
and Part.Intersects calls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:29:43 -04:00