Commit Graph
5 Commits
Author SHA1 Message Date
aj 1b23ad79f2 perf(fill): reuse part triangulations within an overlap check
After 1b, triangulating both polygons on every pair was the largest
remaining overlap cost (27% of main-thread samples on the corpus job).
PartOverlapChecker now triangulates each part at most once per check,
lazily after the bounding-box gate, and passes the triangles to a new
internal Collision.HasOverlap overload that runs the unchanged
OverlapRegions body. Triangles are only read by clipping and hole
subtraction, so reuse gives identical verdicts.

Verification:
- 49,000 seeded decisions with reused triangles match LegacyCollision;
  triangles stay bit-identical to a fresh triangulation afterwards.
- Debug PolygonTriangulations: 246 -> 40 and 64 -> 36 per grid check;
  sharing triangles per Program instead fails 23 tests.
- Corpus job (169 parts, --engines Default --parallel 1): median
  13,398 -> 12,702 ms over 4+4 alternating runs vs 1b, identical
  outcomes; serialized layout byte-identical to the base.

Also records the Follow-up B' (Slices 1a, 1b, 2a) measurements in
docs/performance/fill-performance.md.
2026-09-27 13:49:46 -04:00
aj a27290a29c perf(fill): prepare overlap polygons once per check
Both HasOverlappingParts loops rebuilt each part's polygon from its
Program on every pair. PartOverlapChecker prepares each distinct Program
(reference identity) once and each part's world polygon once per call,
then uses the overlap-only Collision.HasOverlap. Loop order, bounding-box
prefilter, early exit and returned indices are unchanged; Part.Intersects
shares the material/polygon recipe and still returns crossing points.

Verification:
- Frozen LegacyPartOverlap differential (original Intersects and both
  loops): verdicts, indices and world polygons bit-identical across fill
  grids, patterns, touching/epsilon gaps, scribe/rapid/empty programs.
- Debug OverlapPolygonPreparations: 246 -> 1 and 64 -> 2 per check.
- Corpus job (169 parts, --engines Default --parallel 1, with 1a):
  median 18,885 -> 13,464 ms over 4+4 alternating runs, identical
  outcomes; serialized layout byte-identical to the base.
2026-09-27 13:34:04 -04:00
aj 0df2587cf2 perf(fill): reuse offset geometry for translated copies
FillLinear re-prepared offset perimeter geometry (ConvertProgram ->
ShapeProfile -> OffsetOutward) for every part it measured, although
tiled copies share one Program and differ only by Location. A CPU
profile of a 169-part Default job put 62% of wall time there.

Prepare each distinct Program (reference identity) once per public
Fill/FillRow call in local frame, then clone and translate for each
location. The cache is created per call and passed down privately
because FillHelpers.FillPattern calls Fill concurrently on one
instance. PartGeometry gains a local-frame Program overload that the
Part overload now delegates to.

Evaluation order, lazy preparation, fallbacks and tiling are
unchanged. Differential tests against a frozen copy of the previous
FillLinear check bitwise equality, including concurrent calls; Debug
work tests pin preparation counts. With the thread pool capped at one
worker, before/after whole-job layouts are byte-identical. The
Default corpus job median drops from 40,715 to 18,810 ms.
2026-09-26 20:24:34 -04:00
aj 1d1c60daa7 docs(fill): record combined initial-batch acceptance 2026-09-26 11:25:17 -04:00
aj 1e8e532063 perf(fill): skip feature extraction without an angle model 2026-09-26 10:02:15 -04:00