commit f41c5605523abb2ccce6211ce1baf7825f04356d from: ale date: Sun Aug 16 19:47:29 2026 UTC Make INVESTIGATION_LOG.md a permanent docs page; add direnv support INVESTIGATION_LOG.md moves to docs/src/investigation-log.md as a real, committed page instead of being copied there from the repo root at build time — it's the docs site's own content now, not duplicated. README.md's links and docs/make.jl updated accordingly; .gitignore no longer treats that path as a build artifact. Also: .envrc (direnv + nix-direnv, loads flake.nix's devShell — JULIA_PROJECT, julia-bin — automatically on cd, approved via `direnv allow`); examples/.zed/settings.json suppresses julia.lint.missingrefs for that directory, since LanguageServer.jl's static analysis doesn't resolve a package's own self-referential `using` in a standalone script outside its module tree (confirmed a false positive: examples/real_data_demo.jl runs correctly, same environment, all session); .direnv/ (nix-direnv's local build cache) gitignored. commit - 317d1626dbee8cb5405c42c07d09cc722d639559 commit + f41c5605523abb2ccce6211ce1baf7825f04356d blob - /dev/null blob + 3550a30f2de389e537ee40ca5e64a77dc185c79b (mode 644) --- /dev/null +++ .envrc @@ -0,0 +1 @@ +use flake blob - 70e27c1d50ac3da7b1a8b13ae5150a30c547b7f6 blob + 07e07b2301e381aec586ddd124caee2c35572911 --- .gitignore +++ .gitignore @@ -7,11 +7,13 @@ Manifest.toml .vscode/ result result-* +.direnv/ # docs/ — build output and DocumenterVitepress's own Node/Vitepress -# tooling, plus the two source pages generated at build time from -# README.md/INVESTIGATION_LOG.md (see docs/make.jl) — never hand-edited, -# so never committed either. +# tooling, plus docs/src/index.md, generated at build time from +# README.md (see docs/make.jl) and never hand-edited, so never committed +# either. docs/src/investigation-log.md is NOT generated — it's a real, +# permanent page — so it's deliberately not listed here. docs/build/ docs/node_modules/ docs/package-lock.json @@ -20,4 +22,3 @@ docs/.vitepress/dist/ docs/src/.vitepress/cache/ docs/src/.vitepress/dist/ docs/src/index.md -docs/src/investigation-log.md blob - 12d9eb0e2b6e99e1bc4c2b3c654d6b6d34c99d6b (mode 644) blob + /dev/null --- INVESTIGATION_LOG.md +++ /dev/null @@ -1,239 +0,0 @@ -# Investigation log - -Chronological record of what real-data testing found and how each finding -was diagnosed — kept separate from `README.md` so the README stays a -focused reference rather than a narrative. Cross-referenced from -`README.md`'s Known limitations section and from the relevant docstrings. - -## Real ZTF data surfaced four bugs no synthetic test or code review caught - -All fixed, all with regression tests added afterward: - -- An empty cross-match result crashing table construction. -- `crossmatch_catalog(...; :skybot)` silently returning zero matches on - every real call because Julian Date epochs (~2.4e6) print in scientific - notation by default, which SkyBoT's API rejects as an empty epoch. -- `zogy_subtract` never background-subtracted its inputs. A raw FFT's DC - term is a *sum*, not a mean, over the whole image (>10^6 pixels here), - so even a modest sky-brightness mismatch between a science frame and - the reference stack (routine — depends on moon phase and airglow, not - the photometric zeropoint `build_reference` already matches) blew up - into a near-constant offset that swamped `S_corr` almost everywhere: - on real data this meant >99.9% of a frame reading "above 6 sigma" and - zero genuine detections surviving. -- Reprojected frames carry a thin, legitimate NaN border wherever a - dithered footprint doesn't fully cover the target grid; left - unsanitized before `zogy_subtract`, a single NaN poisoned its - whole-array `median`/`fft` calls and turned an entire frame's `S_corr` - to NaN. -- `run_pipeline` filled that same border with `0.0` rather than the valid - region's own background level. Negligible for a real ~10^6 px ZTF frame - (~0.05% invalid), but a regression test using a synthetic frame with a - large dither (~13% invalid) showed this biases `zogy_subtract`'s own - background-median computation toward 0, contaminating the whole - difference image — the same bug, just large enough at that scale to - turn from invisible into a wrong `run_pipeline` output. - -## Dataset-selection mistake: the first real-data demo picked an unrecoverable field - -Not a code bug: the first real-data demo used a field/night where the -faintest catalogued asteroid was 1.2 mag *below* that night's own -detection limit. No algorithm can recover a signal that was never above -the noise. The demo now picks a night with a known object bright enough -to serve as actual ground truth (see `examples/real_data_demo.jl`'s -header comment for how that field/night/reference set was chosen). - -## PSF-timing and astrometric-noise bugs behind ZOGY's excess detections - -Investigating why ZOGY produced ~2x more tracklets than the undifferenced -baseline on real ZTF data (274 vs 129 at the time), with `std(S_corr)` -not actually ~1 as the model assumes and 85-97% of the excess detections -landing within 15 px of a bright reference star, traced it to two -compounding issues: - -- `estimate_psf` measured each frame's PSF *before* reprojection, but - `zogy_subtract` differenced the frame *after*. Reprojection's - interpolation subtly reshapes a PSF, and that mismatch left systematic - subtraction residuals at every bright star. Fixed by measuring the PSF - (and each frame's own detected sources) on the same post-reprojection - pixels `zogy_subtract` actually uses, and feeding those sources through - as `n_sources`/`r_sources` so `zogy_subtract`'s astrometric-noise term - (`V_ast`) can price in residual misregistration instead of ignoring it. -- The reprojected border's NaN was filled with `0.0` rather than the - valid region's own background level (see the bug list above). - -On real data this brought `std(S_corr)` from ~2.0-2.6 down to ~1.1-1.2 in -4 of 5 test frames. The 5th stayed elevated and produced *more* bright -sources than before rather than fewer — investigated separately below. - -## Root-causing the anomalous 5th frame: a likely passing cloud - -GAIN, SEEING, MAGZP, whole-frame background level, saturated-pixel count, -and airmass were all checked directly against the other four frames and -none stood out. Differencing this frame's *raw* pixels against another -frame's, bypassing all ZOGY machinery, found the actual cause: ~1750 -pixels with a significant, one-sided excess (no matching negative lobe, -so not a registration/dipole artifact), ~98% confined to the bottom ~30% -of the detector. A control pair of two normal frames, differenced the -same way, found only ~80-110 such pixels, distributed in proportion to -where the frame's actual stars are — not concentrated in one region. The -whole-frame background stayed flat across that region (no gradient), -ruling out amplifier glow or vignetting. - -This pattern — real stars showing localized excess halos in one part of -the frame, flat background, no positional offset — is the signature of -thin cloud scattering starlight during the 30 s exposure, over only part -of the field, not a code or calibration bug. It also happens to be the -one frame among the five with an invalid (negative) `MOONILLF` value in -its own archived metadata, consistent with degraded weather telemetry at -that time. Not independently confirmed against an all-sky camera or cloud -sensor, but every alternative explanation checked was ruled out by direct -measurement rather than left as a guess. - -## Quality gate: MAD looked more rigorous, but was blind to the real anomaly - -A per-frame quality gate (`run_pipeline`'s `quality_max_std`) was added to -catch frames like the one above. First attempt used a MAD-based -(median-absolute-deviation) robust spread instead of plain -`Statistics.std`, specifically to avoid a real bright source skewing -`std` in a small *synthetic test* image (confirmed: at 100x100 px, a -genuine ~60 sigma injected source pushed a clean frame's std to ~4). - -Against real data this backfired: the actual anomalous frame's excess -showed up as ~232 point-like bright residuals, not a bulk noise increase -— exactly the kind of minority-of-pixels outlier MAD is *designed* to -ignore. Its MAD measured ~0.997, indistinguishable from the good frames' -~0.98-1.10, while its plain `std` (2.025) stood clearly apart from the -good frames' (1.10-1.18). Reverted to plain `std`; fixed the synthetic -test's false positive by using a larger image (600x600) instead of -changing the statistic, since on real ~10^6 px data a genuine source's -few dozen pixels are too small a fraction to skew `std` regardless. - -Confirmed against real data: the gate (default `quality_max_std=1.5`) -correctly excludes the anomalous frame (`frame_std=2.025 > 1.5`), and -both known SkyBoT objects are still recovered afterward. - -## The quality gate's combinatorial side effect on tracklet count - -`link_candidates` requires every frame to match by default -(`min_frames`). A gated-out frame contributes zero detections, so once -the quality gate correctly excludes one, no tracklet is reachable at all -unless `min_frames` is lowered to account for it — -`examples/real_data_demo.jl` uses `length(fits_paths) - 1`. - -That one-frame relaxation, independent of anything about data quality, -roughly doubled the tracklet count again (334 → 667) purely through -looser matching combinatorics: requiring 4-of-4 frames to match is a -weaker constraint than requiring 5-of-5, so more spurious coincidental -matches pass regardless of how clean the underlying data is. The gate -itself is confirmed working (see above), but this means the real-data -tracklet-count comparison in `README.md` is *not* a clean read on ZOGY's -noise properties — the `min_frames` change confounds it. A clean -before/after comparison would need to hold `min_frames` fixed some other -way (e.g. by literally excluding the bad frame from the input list rather -than gating it internally), not attempted here. - -## `plate_solve` validated end-to-end against the live service - -nova.astrometry.net's own API docs require a `Referer: -https://nova.astrometry.net/api/login` header on programmatic file -downloads, as an anti-scraper-bot check — `plate_solve` didn't set it. -Added it to every GET request in `src/platesolve.jl` (submission-status -poll, job-status poll, WCS fetch); also switched the base URLs from -`http://` to `https://` to match the documented endpoints. - -With a real API key: a synthetic image containing one fabricated point -source in noise (no genuine star pattern) correctly failed to solve, -astrometry.net reporting `status: "failure"` rather than returning a -wrong answer — the expected outcome, since plate-solving works by -matching real star-pattern asterisms against a sky index, and there was -no real pattern to match. A real ZTF frame (`data/real/science/`, its own -existing WCS ignored) solved successfully: recovered -`CRVAL ≈ (36.3997, 2.0038)`, consistent with that field's known centre -(~36.5, ~2.1). One upload attempt hit a transient `503` from the -service; a retry succeeded, so `plate_solve` callers should be prepared -to retry on transient server errors rather than assume a single failed -upload means the service is unusable. - -## `detect_sources`'s flux has been silently wrong since the beginning: a transposed aperture - -Building `find_variable_sources` required trusting `detect_sources`'s -`flux` for the first time in a *quantitative* way — every earlier use -(`link_candidates`, tracklet building, WCS calibration, cross-matching) -only depends on its `x`/`y` positions. That new dependency surfaced a bug -that had been present, silently, since `detect_sources` was first written: -its measured flux was wrong by orders of magnitude for almost any real -source. - -Root cause: `Photometry.jl` (v0.9.8) is internally inconsistent between -its own two pieces. `PeakMesh`'s `extract_sources` reports positions in -the standard Cartesian sense — `to_nt(ci) = (x=ci[2], y=ci[1], ...)`, i.e. -`x` is the array's *second* dimension (column), `y` its *first* (row). -`CircularAperture`, in the same package, does the opposite internally — -`bounds`/`overlap` treat its own `.x` field as indexing the array's -*first* dimension and `.y` the second. `detect_sources` built -`CircularAperture(row.x, row.y, aperture_radius)` directly from -`PeakMesh`'s output, so every aperture was centred at the transposed -pixel — correct only when a source's row and column indices happened to -coincide (the diagonal), or invisibly wrong on a square canvas at a -generic position, or (on a non-square canvas, or once truly out of the -transposed array's bounds) landing on pure background instead. - -Found by comparing `detect_sources`'s measured flux against the closed-form -enclosed-energy integral for an isolated, noise-free synthetic Gaussian on -a non-square canvas at an asymmetric position: measured flux came back -~140x too small. Swapping the aperture's constructor arguments -(`CircularAperture(row.y, row.x, aperture_radius)`) brought it to within -~1.4% of the analytic value — geometric quantization error, not a -remaining bug. `light_curve` (`src/rotation.jl`) was checked the same way -and is *not* affected: it deliberately never permutes its image (see its -own docstring), and `WCS.world_to_pix`'s returned `(x, y)` already matches -raw FITS `(NAXIS1, NAXIS2)` order — which happens to be exactly what -`CircularAperture` expects internally, by coincidence of two conventions -cancelling out rather than by any intentional match. - -This had zero effect on every result validated so far in this project — -`run_pipeline`'s real-data tracklet counts (133 baseline / 667 ZOGY, both -recovering the same 2 known objects) depend only on detected *positions*, -never on `flux` — but it fully invalidated the first real-data -measurements made *while building* `find_variable_sources` earlier the -same session, before this was found: a claimed 29.4% median -`flux_err/flux` (actual, post-fix: **2.56%**), a claimed "only 2 of 120 -stars clear a 10% S/N floor" (actual: **119 of 119**), and a claimed -8-40% ensemble-vs-`MAGZP` mismatch explained by low-S/N selection bias -(the mismatch is real, but stable at 9-13% with or without an S/N cut — -not a selection effect at all; see below). Every constant and docstring -claim built on those numbers was rewritten against the corrected -measurements rather than left standing. - -## What was actually driving the ~10 percent photometric-scale mismatch, once flux was measured correctly - -With flux measured correctly, `photometric_scale`'s ensemble ratio against -each frame's own `MAGZP` zeropoint still disagreed by 9-13% — but now -*independent* of the S/N cut, ruling out selection bias as the cause. -Checked against each frame's `SEEING` header value (1.805-2.009 px across -the 5 real frames): the frames with better (smaller) seeing than frame 1 -all showed `photometric_scale < 1`, consistent with the standard "aperture -correction" effect — a fixed-radius aperture (`aperture_radius`, default 3 -px, comparable to ZTF's own seeing) encloses a larger fraction of a star's -total flux when the PSF is more concentrated. `photometric_scale`'s -ensemble-differential approach corrects for this automatically (it only -needs frame-to-frame consistency, not a causal model), so no code change -was needed here — only the docstring's explanation, which had cited the -now-debunked selection-bias story. - -## `find_variable_sources`'s chi-squared test has a real, measured false-positive floor on real data - -Even with correct flux and photometric normalization, real ZTF data's -reduced-chi2 distribution over 119 matched stationary stars has a heavy -tail: 23% exceed a threshold of 3 (the textbook-reasonable default), -13% still exceed 20. The likely cause: `detect_sources` positions each -detection at `PeakMesh`'s integer peak pixel, not a sub-pixel centroid, so -a 1-pixel jitter between frames — from noise, or a slightly different PSF -realization — against a small aperture (3 px, close to the PSF's own -scale) produces a real, but spurious, flux swing frame to frame. Raising -`chi2_threshold`'s default to 50 cuts the false-positive rate to 8% on -this dataset — better, but not clean, and documented as such in -`find_variable_sources`'s own docstring rather than presented as solved. -The actual fix (forced, sub-pixel-centroided photometry instead of -peak-position aperture photometry) is a larger change, not attempted here. blob - 7ba634c95babb8347b5a571869a28e2e80bbf93c blob + ffb9f60bdf56e658cba68c13f6d1f85f9e52e079 --- README.md +++ README.md @@ -73,7 +73,7 @@ are implemented, tested against synthetic data, and †`find_variable_sources`'s photometric normalization, S/N floor, and `chi2_threshold` default, and for `fit_moffat_psf`'s recovered PSF width — calibrated directly against real ZTF data (see -[`INVESTIGATION_LOG.md`](INVESTIGATION_LOG.md), including a real bug in +[`INVESTIGATION_LOG.md`](docs/src/investigation-log.md), including a real bug in `detect_sources`'s own flux measurement this calibration work found and fixed). Not yet exercised end to end, via `search_field`, against a real field with an independently-confirmed variable star as ground truth — the @@ -89,7 +89,7 @@ here are bright enough that the baseline already recov so this dataset doesn't exercise ZOGY's actual advantage (recovering objects below a single frame's noise floor). The raw tracklet-count gap is not a clean read on ZOGY's noise properties; see -[`INVESTIGATION_LOG.md`](INVESTIGATION_LOG.md) for why, and for the full +[`INVESTIGATION_LOG.md`](docs/src/investigation-log.md) for why, and for the full record of every real bug this project's real-data testing surfaced (five so far, all fixed with regression tests) and how each was diagnosed. @@ -110,7 +110,7 @@ so far, all fixed with regression tests) and how each - **A quality-gated frame silently tightens `link_candidates`.** `run_pipeline`'s `quality_max_std` (default `1.5`) excludes a frame whose `S_corr` spread is too high (confirmed against real data — see - [`INVESTIGATION_LOG.md`](INVESTIGATION_LOG.md)), but a gated frame + [`INVESTIGATION_LOG.md`](docs/src/investigation-log.md)), but a gated frame contributes zero detections, and `link_candidates` requires every frame to match by default. Pass a lower `min_frames` (e.g. `length(fits_paths) - 1`) when using the ZOGY path, or no tracklet will @@ -122,12 +122,12 @@ so far, all fixed with regression tests) and how each look like genuine variability; on real ZTF data even a generous `chi2_threshold` still flags several times more stars than the true stellar variable fraction (see `find_variable_sources`'s docstring and - [`INVESTIGATION_LOG.md`](INVESTIGATION_LOG.md) for the measured rate). + [`INVESTIGATION_LOG.md`](docs/src/investigation-log.md) for the measured rate). Treat a candidate as needing independent confirmation (a catalog match or a recovered period), not as self-evidently real. - **`crossmatch_catalog(...; :vsx)`/`(...; :simbad)` query one candidate at a time.** Migrated off the CDS X-Match service (extended, total - outages — see [`INVESTIGATION_LOG.md`](INVESTIGATION_LOG.md)) to direct + outages — see [`INVESTIGATION_LOG.md`](docs/src/investigation-log.md)) to direct SIMBAD/VizieR TAP queries, which don't offer X-Match's single-batched-request shape; a large candidate list means that many requests. `:skybot` is unaffected (a different service, always queried this way). @@ -155,6 +155,8 @@ target should barely move between them, unlike the ori epochs): ```julia +using AsteroidPipeline + times, flux, flux_err = light_curve(fits_paths, ra, dec) result = recover_rotation_period(times, flux; minimum_period=0.02, maximum_period=1.0) result.period, result.false_alarm_probability @@ -173,13 +175,15 @@ For a frame with no WCS already in its header, and a registration): ```julia +using AsteroidPipeline + run_pipeline(fits_paths; reference=reference, plate_solve_api_key=key) ``` or directly: `plate_solve(fits_path; api_key=key)`. This is a live network round trip — upload, then poll until the frame solves — so it is slow and requires connectivity. Validated against the real service: see -[`INVESTIGATION_LOG.md`](INVESTIGATION_LOG.md). +[`INVESTIGATION_LOG.md`](docs/src/investigation-log.md). ## Using real IASC campaign data @@ -222,12 +226,14 @@ nix develop ## Documentation The full API reference (every exported function's docstring, organized by -pipeline stage) plus this README and `INVESTIGATION_LOG.md` are built into +pipeline stage) plus this README and the investigation log are built into a static site with [Documenter.jl](https://github.com/JuliaDocs/Documenter.jl) and [DocumenterVitepress.jl](https://github.com/LuxDL/DocumenterVitepress.jl). -`docs/make.jl` copies `README.md`/`INVESTIGATION_LOG.md` into `docs/src/` -at build time (not committed — see `.gitignore`), so the site can't drift -out of sync with them. +`docs/src/investigation-log.md` (linked throughout this README as +`INVESTIGATION_LOG.md`) is a real, permanent page there — the docs site's +own content, not copied in from elsewhere. `docs/make.jl` does copy +`README.md` itself into `docs/src/index.md` at build time (not committed — +see `.gitignore`), so the site's home page can't drift out of sync with it. To build locally: blob - 72455561ca3ff6ff2568cf6fcaa5aad5d5af895d blob + 02573bb4fc8847b0a41cf1810c08643dcf32809c --- docs/make.jl +++ docs/make.jl @@ -2,19 +2,18 @@ using AsteroidPipeline using Documenter using DocumenterVitepress -# README.md and INVESTIGATION_LOG.md are the actual source of truth for -# the project's purpose, status, and real-data findings (see -# INVESTIGATION_LOG.md itself for why they're kept separate). Copied here -# at build time, not committed under docs/src (see .gitignore), so the -# docs site can never drift out of sync with them. +# README.md is the actual source of truth for the project's purpose and +# status. Copied here at build time, not committed under docs/src (see +# .gitignore), so the site's home page can never drift out of sync with +# it. INVESTIGATION_LOG.md lives permanently at docs/src/investigation-log.md +# instead (a real, committed page, not generated) — it's the docs site's +# own content, not duplicated from anywhere else. const REPO_ROOT = joinpath(@__DIR__, "..") index_content = read(joinpath(REPO_ROOT, "README.md"), String) -index_content = replace(index_content, "(INVESTIGATION_LOG.md)" => "(investigation-log.md)") +index_content = replace(index_content, "(docs/src/investigation-log.md)" => "(investigation-log.md)") write(joinpath(@__DIR__, "src", "index.md"), index_content) -cp(joinpath(REPO_ROOT, "INVESTIGATION_LOG.md"), joinpath(@__DIR__, "src", "investigation-log.md"); force=true) - DocMeta.setdocmeta!(AsteroidPipeline, :DocTestSetup, :(using AsteroidPipeline); recursive=true) makedocs(; blob - /dev/null blob + 12d9eb0e2b6e99e1bc4c2b3c654d6b6d34c99d6b (mode 644) --- /dev/null +++ docs/src/investigation-log.md @@ -0,0 +1,239 @@ +# Investigation log + +Chronological record of what real-data testing found and how each finding +was diagnosed — kept separate from `README.md` so the README stays a +focused reference rather than a narrative. Cross-referenced from +`README.md`'s Known limitations section and from the relevant docstrings. + +## Real ZTF data surfaced four bugs no synthetic test or code review caught + +All fixed, all with regression tests added afterward: + +- An empty cross-match result crashing table construction. +- `crossmatch_catalog(...; :skybot)` silently returning zero matches on + every real call because Julian Date epochs (~2.4e6) print in scientific + notation by default, which SkyBoT's API rejects as an empty epoch. +- `zogy_subtract` never background-subtracted its inputs. A raw FFT's DC + term is a *sum*, not a mean, over the whole image (>10^6 pixels here), + so even a modest sky-brightness mismatch between a science frame and + the reference stack (routine — depends on moon phase and airglow, not + the photometric zeropoint `build_reference` already matches) blew up + into a near-constant offset that swamped `S_corr` almost everywhere: + on real data this meant >99.9% of a frame reading "above 6 sigma" and + zero genuine detections surviving. +- Reprojected frames carry a thin, legitimate NaN border wherever a + dithered footprint doesn't fully cover the target grid; left + unsanitized before `zogy_subtract`, a single NaN poisoned its + whole-array `median`/`fft` calls and turned an entire frame's `S_corr` + to NaN. +- `run_pipeline` filled that same border with `0.0` rather than the valid + region's own background level. Negligible for a real ~10^6 px ZTF frame + (~0.05% invalid), but a regression test using a synthetic frame with a + large dither (~13% invalid) showed this biases `zogy_subtract`'s own + background-median computation toward 0, contaminating the whole + difference image — the same bug, just large enough at that scale to + turn from invisible into a wrong `run_pipeline` output. + +## Dataset-selection mistake: the first real-data demo picked an unrecoverable field + +Not a code bug: the first real-data demo used a field/night where the +faintest catalogued asteroid was 1.2 mag *below* that night's own +detection limit. No algorithm can recover a signal that was never above +the noise. The demo now picks a night with a known object bright enough +to serve as actual ground truth (see `examples/real_data_demo.jl`'s +header comment for how that field/night/reference set was chosen). + +## PSF-timing and astrometric-noise bugs behind ZOGY's excess detections + +Investigating why ZOGY produced ~2x more tracklets than the undifferenced +baseline on real ZTF data (274 vs 129 at the time), with `std(S_corr)` +not actually ~1 as the model assumes and 85-97% of the excess detections +landing within 15 px of a bright reference star, traced it to two +compounding issues: + +- `estimate_psf` measured each frame's PSF *before* reprojection, but + `zogy_subtract` differenced the frame *after*. Reprojection's + interpolation subtly reshapes a PSF, and that mismatch left systematic + subtraction residuals at every bright star. Fixed by measuring the PSF + (and each frame's own detected sources) on the same post-reprojection + pixels `zogy_subtract` actually uses, and feeding those sources through + as `n_sources`/`r_sources` so `zogy_subtract`'s astrometric-noise term + (`V_ast`) can price in residual misregistration instead of ignoring it. +- The reprojected border's NaN was filled with `0.0` rather than the + valid region's own background level (see the bug list above). + +On real data this brought `std(S_corr)` from ~2.0-2.6 down to ~1.1-1.2 in +4 of 5 test frames. The 5th stayed elevated and produced *more* bright +sources than before rather than fewer — investigated separately below. + +## Root-causing the anomalous 5th frame: a likely passing cloud + +GAIN, SEEING, MAGZP, whole-frame background level, saturated-pixel count, +and airmass were all checked directly against the other four frames and +none stood out. Differencing this frame's *raw* pixels against another +frame's, bypassing all ZOGY machinery, found the actual cause: ~1750 +pixels with a significant, one-sided excess (no matching negative lobe, +so not a registration/dipole artifact), ~98% confined to the bottom ~30% +of the detector. A control pair of two normal frames, differenced the +same way, found only ~80-110 such pixels, distributed in proportion to +where the frame's actual stars are — not concentrated in one region. The +whole-frame background stayed flat across that region (no gradient), +ruling out amplifier glow or vignetting. + +This pattern — real stars showing localized excess halos in one part of +the frame, flat background, no positional offset — is the signature of +thin cloud scattering starlight during the 30 s exposure, over only part +of the field, not a code or calibration bug. It also happens to be the +one frame among the five with an invalid (negative) `MOONILLF` value in +its own archived metadata, consistent with degraded weather telemetry at +that time. Not independently confirmed against an all-sky camera or cloud +sensor, but every alternative explanation checked was ruled out by direct +measurement rather than left as a guess. + +## Quality gate: MAD looked more rigorous, but was blind to the real anomaly + +A per-frame quality gate (`run_pipeline`'s `quality_max_std`) was added to +catch frames like the one above. First attempt used a MAD-based +(median-absolute-deviation) robust spread instead of plain +`Statistics.std`, specifically to avoid a real bright source skewing +`std` in a small *synthetic test* image (confirmed: at 100x100 px, a +genuine ~60 sigma injected source pushed a clean frame's std to ~4). + +Against real data this backfired: the actual anomalous frame's excess +showed up as ~232 point-like bright residuals, not a bulk noise increase +— exactly the kind of minority-of-pixels outlier MAD is *designed* to +ignore. Its MAD measured ~0.997, indistinguishable from the good frames' +~0.98-1.10, while its plain `std` (2.025) stood clearly apart from the +good frames' (1.10-1.18). Reverted to plain `std`; fixed the synthetic +test's false positive by using a larger image (600x600) instead of +changing the statistic, since on real ~10^6 px data a genuine source's +few dozen pixels are too small a fraction to skew `std` regardless. + +Confirmed against real data: the gate (default `quality_max_std=1.5`) +correctly excludes the anomalous frame (`frame_std=2.025 > 1.5`), and +both known SkyBoT objects are still recovered afterward. + +## The quality gate's combinatorial side effect on tracklet count + +`link_candidates` requires every frame to match by default +(`min_frames`). A gated-out frame contributes zero detections, so once +the quality gate correctly excludes one, no tracklet is reachable at all +unless `min_frames` is lowered to account for it — +`examples/real_data_demo.jl` uses `length(fits_paths) - 1`. + +That one-frame relaxation, independent of anything about data quality, +roughly doubled the tracklet count again (334 → 667) purely through +looser matching combinatorics: requiring 4-of-4 frames to match is a +weaker constraint than requiring 5-of-5, so more spurious coincidental +matches pass regardless of how clean the underlying data is. The gate +itself is confirmed working (see above), but this means the real-data +tracklet-count comparison in `README.md` is *not* a clean read on ZOGY's +noise properties — the `min_frames` change confounds it. A clean +before/after comparison would need to hold `min_frames` fixed some other +way (e.g. by literally excluding the bad frame from the input list rather +than gating it internally), not attempted here. + +## `plate_solve` validated end-to-end against the live service + +nova.astrometry.net's own API docs require a `Referer: +https://nova.astrometry.net/api/login` header on programmatic file +downloads, as an anti-scraper-bot check — `plate_solve` didn't set it. +Added it to every GET request in `src/platesolve.jl` (submission-status +poll, job-status poll, WCS fetch); also switched the base URLs from +`http://` to `https://` to match the documented endpoints. + +With a real API key: a synthetic image containing one fabricated point +source in noise (no genuine star pattern) correctly failed to solve, +astrometry.net reporting `status: "failure"` rather than returning a +wrong answer — the expected outcome, since plate-solving works by +matching real star-pattern asterisms against a sky index, and there was +no real pattern to match. A real ZTF frame (`data/real/science/`, its own +existing WCS ignored) solved successfully: recovered +`CRVAL ≈ (36.3997, 2.0038)`, consistent with that field's known centre +(~36.5, ~2.1). One upload attempt hit a transient `503` from the +service; a retry succeeded, so `plate_solve` callers should be prepared +to retry on transient server errors rather than assume a single failed +upload means the service is unusable. + +## `detect_sources`'s flux has been silently wrong since the beginning: a transposed aperture + +Building `find_variable_sources` required trusting `detect_sources`'s +`flux` for the first time in a *quantitative* way — every earlier use +(`link_candidates`, tracklet building, WCS calibration, cross-matching) +only depends on its `x`/`y` positions. That new dependency surfaced a bug +that had been present, silently, since `detect_sources` was first written: +its measured flux was wrong by orders of magnitude for almost any real +source. + +Root cause: `Photometry.jl` (v0.9.8) is internally inconsistent between +its own two pieces. `PeakMesh`'s `extract_sources` reports positions in +the standard Cartesian sense — `to_nt(ci) = (x=ci[2], y=ci[1], ...)`, i.e. +`x` is the array's *second* dimension (column), `y` its *first* (row). +`CircularAperture`, in the same package, does the opposite internally — +`bounds`/`overlap` treat its own `.x` field as indexing the array's +*first* dimension and `.y` the second. `detect_sources` built +`CircularAperture(row.x, row.y, aperture_radius)` directly from +`PeakMesh`'s output, so every aperture was centred at the transposed +pixel — correct only when a source's row and column indices happened to +coincide (the diagonal), or invisibly wrong on a square canvas at a +generic position, or (on a non-square canvas, or once truly out of the +transposed array's bounds) landing on pure background instead. + +Found by comparing `detect_sources`'s measured flux against the closed-form +enclosed-energy integral for an isolated, noise-free synthetic Gaussian on +a non-square canvas at an asymmetric position: measured flux came back +~140x too small. Swapping the aperture's constructor arguments +(`CircularAperture(row.y, row.x, aperture_radius)`) brought it to within +~1.4% of the analytic value — geometric quantization error, not a +remaining bug. `light_curve` (`src/rotation.jl`) was checked the same way +and is *not* affected: it deliberately never permutes its image (see its +own docstring), and `WCS.world_to_pix`'s returned `(x, y)` already matches +raw FITS `(NAXIS1, NAXIS2)` order — which happens to be exactly what +`CircularAperture` expects internally, by coincidence of two conventions +cancelling out rather than by any intentional match. + +This had zero effect on every result validated so far in this project — +`run_pipeline`'s real-data tracklet counts (133 baseline / 667 ZOGY, both +recovering the same 2 known objects) depend only on detected *positions*, +never on `flux` — but it fully invalidated the first real-data +measurements made *while building* `find_variable_sources` earlier the +same session, before this was found: a claimed 29.4% median +`flux_err/flux` (actual, post-fix: **2.56%**), a claimed "only 2 of 120 +stars clear a 10% S/N floor" (actual: **119 of 119**), and a claimed +8-40% ensemble-vs-`MAGZP` mismatch explained by low-S/N selection bias +(the mismatch is real, but stable at 9-13% with or without an S/N cut — +not a selection effect at all; see below). Every constant and docstring +claim built on those numbers was rewritten against the corrected +measurements rather than left standing. + +## What was actually driving the ~10 percent photometric-scale mismatch, once flux was measured correctly + +With flux measured correctly, `photometric_scale`'s ensemble ratio against +each frame's own `MAGZP` zeropoint still disagreed by 9-13% — but now +*independent* of the S/N cut, ruling out selection bias as the cause. +Checked against each frame's `SEEING` header value (1.805-2.009 px across +the 5 real frames): the frames with better (smaller) seeing than frame 1 +all showed `photometric_scale < 1`, consistent with the standard "aperture +correction" effect — a fixed-radius aperture (`aperture_radius`, default 3 +px, comparable to ZTF's own seeing) encloses a larger fraction of a star's +total flux when the PSF is more concentrated. `photometric_scale`'s +ensemble-differential approach corrects for this automatically (it only +needs frame-to-frame consistency, not a causal model), so no code change +was needed here — only the docstring's explanation, which had cited the +now-debunked selection-bias story. + +## `find_variable_sources`'s chi-squared test has a real, measured false-positive floor on real data + +Even with correct flux and photometric normalization, real ZTF data's +reduced-chi2 distribution over 119 matched stationary stars has a heavy +tail: 23% exceed a threshold of 3 (the textbook-reasonable default), +13% still exceed 20. The likely cause: `detect_sources` positions each +detection at `PeakMesh`'s integer peak pixel, not a sub-pixel centroid, so +a 1-pixel jitter between frames — from noise, or a slightly different PSF +realization — against a small aperture (3 px, close to the PSF's own +scale) produces a real, but spurious, flux swing frame to frame. Raising +`chi2_threshold`'s default to 50 cuts the false-positive rate to 8% on +this dataset — better, but not clean, and documented as such in +`find_variable_sources`'s own docstring rather than presented as solved. +The actual fix (forced, sub-pixel-centroided photometry instead of +peak-position aperture photometry) is a larger change, not attempted here. blob - /dev/null blob + 98e288dec0b717e3d0d0d6706060fee5ac867c79 (mode 644) --- /dev/null +++ examples/.zed/settings.json @@ -0,0 +1,9 @@ +{ + "lsp": { + "julia": { + "settings": { + "julia.lint.missingrefs": "none" + } + } + } +}