commit - 6360d2173215c81e8cf12c81a99bb96eb23dce9b
commit + 5dd14c0e4f1d597bb33e97a415614e28c99f1ec4
blob - 73597de260b27d8f7525c16d71419d0145c78157
blob + 55bdce7f712a94eb3c5539c0e07dd5206cac5f87
--- docs/src/design-refinements.md
+++ docs/src/design-refinements.md
a plain FITS header string (just a `String`, no pointers, serializes
safely) — send that instead, and have each worker rebuild its own local
`WCSTransform` via [`load_wcs`](@ref) before calling
-`Reproject.reproject`. Confirmed on the same real 30-frame field-451
-reference set, 8 worker processes: no crash, and a real 331.8s vs. the
-~720s sequential baseline above — about **2.2x**, not full 8x
-core-count scaling, since each worker still re-parses its own WCS header
-per call and `pmap`'s own scheduling/serialization isn't free. Landed as
+`Reproject.reproject`. Confirmed end to end on `build_reference` itself
+(not just the technique standalone), on the same real 30-frame field-451
+reference set, 8 worker processes, sequential and distributed measured
+back to back with identical output: no crash, 295.92s vs. 951.31s
+sequential — a real **3.21x**, not full 8x core-count scaling, since
+each worker still re-parses its own WCS header per call and `pmap`'s own
+scheduling/serialization isn't free. Landed as
`build_reference`'s `workers` keyword (opt-in; the caller supplies
already-running worker processes, since spawning and managing a process
pool is an environment concern, not something a data-processing function
blob - 61d7ae6f5b300a3eff0b37be6fc96410ccd97205
blob + 40d4aea1d034d4aab46e41e095b92745b2f76df4
--- src/reference.jl
+++ src/reference.jl
is converted to a plain FITS header string (`WCS.to_header`, no pointers,
serializes safely) before being sent, and each worker reconstructs its
own local `WCSTransform` from that string (`load_wcs`) before calling
-`Reproject.reproject` — confirmed on real data (real 30-frame ZTF
+`Reproject.reproject` — confirmed end to end, on this actual function
+(not just the technique in isolation), on real data (real 30-frame ZTF
field-451 reference set, 8 worker processes) to run without crashing,
-in 331.8s vs. an ~720s sequential baseline (the ~24s/frame figure above)
-— about 2.2x, not full 8x core-count scaling, since each worker still
-does its own `load_wcs` parsing per call and `pmap`'s own scheduling and
+in 295.92s vs. 951.31s sequential, both measured back to back with
+identical output (`image`/`sigma`/`mask` all exactly equal) — a real
+3.21x, not full 8x core-count scaling, since each worker still does its
+own `load_wcs` parsing per call and `pmap`'s own scheduling and
serialization overhead isn't free. `workers` must be pre-existing worker
process ids (e.g. from `Distributed.addprocs`), each of which the caller
must have already loaded this package on (`@everywhere using