commit 5dd14c0e4f1d597bb33e97a415614e28c99f1ec4 from: ale date: Wed Aug 19 02:45:37 2026 UTC Update build_reference's benchmark numbers with the real, integrated result The earlier 331.8s/2.2x came from a standalone script that reimplemented the WCS-header-string technique by hand, before build_reference's own workers keyword existed. Re-measured end to end on the actual function, same real 30-frame field-451 set, sequential and distributed back to back: 951.31s vs. 295.92s, 3.21x — better than the earlier estimate, and output (image/sigma/mask) confirmed exactly identical between the two paths. commit - 6360d2173215c81e8cf12c81a99bb96eb23dce9b commit + 5dd14c0e4f1d597bb33e97a415614e28c99f1ec4 blob - 73597de260b27d8f7525c16d71419d0145c78157 blob + 55bdce7f712a94eb3c5539c0e07dd5206cac5f87 --- docs/src/design-refinements.md +++ docs/src/design-refinements.md @@ -187,11 +187,13 @@ process. The fix: never let a `WCSTransform` cross the a plain FITS header string (just a `String`, no pointers, serializes safely) — send that instead, and have each worker rebuild its own local `WCSTransform` via [`load_wcs`](@ref) before calling -`Reproject.reproject`. Confirmed on the same real 30-frame field-451 -reference set, 8 worker processes: no crash, and a real 331.8s vs. the -~720s sequential baseline above — about **2.2x**, not full 8x -core-count scaling, since each worker still re-parses its own WCS header -per call and `pmap`'s own scheduling/serialization isn't free. Landed as +`Reproject.reproject`. Confirmed end to end on `build_reference` itself +(not just the technique standalone), on the same real 30-frame field-451 +reference set, 8 worker processes, sequential and distributed measured +back to back with identical output: no crash, 295.92s vs. 951.31s +sequential — a real **3.21x**, not full 8x core-count scaling, since +each worker still re-parses its own WCS header per call and `pmap`'s own +scheduling/serialization isn't free. Landed as `build_reference`'s `workers` keyword (opt-in; the caller supplies already-running worker processes, since spawning and managing a process pool is an environment concern, not something a data-processing function blob - 61d7ae6f5b300a3eff0b37be6fc96410ccd97205 blob + 40d4aea1d034d4aab46e41e095b92745b2f76df4 --- src/reference.jl +++ src/reference.jl @@ -40,11 +40,13 @@ the main process). Passing `workers` here avoids that: is converted to a plain FITS header string (`WCS.to_header`, no pointers, serializes safely) before being sent, and each worker reconstructs its own local `WCSTransform` from that string (`load_wcs`) before calling -`Reproject.reproject` — confirmed on real data (real 30-frame ZTF +`Reproject.reproject` — confirmed end to end, on this actual function +(not just the technique in isolation), on real data (real 30-frame ZTF field-451 reference set, 8 worker processes) to run without crashing, -in 331.8s vs. an ~720s sequential baseline (the ~24s/frame figure above) -— about 2.2x, not full 8x core-count scaling, since each worker still -does its own `load_wcs` parsing per call and `pmap`'s own scheduling and +in 295.92s vs. 951.31s sequential, both measured back to back with +identical output (`image`/`sigma`/`mask` all exactly equal) — a real +3.21x, not full 8x core-count scaling, since each worker still does its +own `load_wcs` parsing per call and `pmap`'s own scheduling and serialization overhead isn't free. `workers` must be pre-existing worker process ids (e.g. from `Distributed.addprocs`), each of which the caller must have already loaded this package on (`@everywhere using