Independent Research · Machine Learning × Astronomy
Do Lens Detectors Need Look-Alike Negatives?
A controlled experiment on what actually fixes false positives in strong gravitational-lens detectors — and a preregistered follow-up that was ended by its own kill-switch.
Summary
Strong gravitational-lens searches drown in false positives, and a popular fix is to retrain the detector on hard negative examples that look like lenses. This study asked whether it is the look-alike morphology that helps, or simply having more matched negatives at all. Under a controlled protocol, ordinary galaxies matched on brightness and redshift reduced false positives about as much as the classic spiral look-alikes did.
Status & scope
Manuscript in preparation. No lenses were discovered in this work, and the detector is not deployment-ready. The ten "possible lens" flags from blind labeling are one person's visual triage of small image cutouts — not candidates and not discoveries.
Overview
This project is a study of how machine-learning detectors of strong gravitational lenses should be trained — not an attempt to build a better detector. In one sentence: a controlled experiment showed that a popular fix for lens-detector false positives works mainly because it adds matched training data, not because of the specific look-alike morphology — followed by a preregistered extension that was terminated by its own kill-switch when it failed to beat the bar it had set for itself.
The work has three chapters: a benchmark detector that looked excellent and wasn't, the controlled experiment that took apart an earlier finding of mine, and a follow-up whose success criterion was written down and frozen before any results existed — and which stopped the project when it wasn't met.
Background
A strong gravitational lens is a rare alignment in which a massive galaxy bends the light of a more distant one into arcs or rings. Lenses are valuable to cosmologists and extremely rare — roughly one in ten thousand galaxies — so modern searches are automated, and automated searches drown in false positives: ordinary galaxies whose shapes mimic lensed arcs.
All data in this project are fully public: image cutouts from the DESI Legacy Imaging Surveys (DR9), citizen-science morphology votes from Galaxy Zoo DECaLS, published lens catalogs (Huang et al. 2020, 2021), redMaPPer SDSS cluster members (Rykoff et al. 2014), which supplied the baseline's training negatives, and SDSS DR17 spectroscopic galaxies. The work was done solo, with no lab and no mentor.
The failure that started it
A detector scoring roughly 0.99 AUC on its own benchmark — apparently excellent — turned out to false-positive heavily on realistic galaxy populations. On a frozen, never-before-scored set of merging galaxies, the baseline flagged 27–57% of them as lens candidates, depending on architecture. Those are means across ten seeds at thresholds frozen before scoring; individual seeds swing much wider, which is its own argument against trusting any single training run. The benchmark number and the real-world behavior were telling two different stories, and the rest of the project came out of taking that gap seriously.
The controlled experiment
An earlier result of mine suggested a fix: training on spiral galaxies, the classic arc look-alike, appeared to repair the false-positive problem. But that comparison changed three things at once — the amount of training data, the class balance, and the training exposure. So the apparent "spiral effect" had never actually been isolated.
The centerpiece of the project is a proper control: 90 training runs (3 architectures × 10 seeds × 3 conditions). Every run of a given architecture and seed started from one hashed initial state, cloned bit-for-bit into all three conditions, and every decision threshold was frozen before any test data was scored. The two intervention arms — S and Q, the comparison the study actually turns on — match each other exactly on training-set size, class balance, and optimizer steps. The baseline arm is deliberately smaller (225 steps against 315) and is reported as-is, because its smaller size is the historical confound the study set out to isolate.
| Condition | Training set | False-positive change |
|---|---|---|
| A — baseline | Unmodified training set | — |
| S — spiral look-alikes | Baseline + 175 spiral galaxies, the classic arc mimic | −40.3 points |
| Q — matched ordinary galaxies | Baseline + 175 ordinary galaxies, matched one-to-one to the spirals on brightness and redshift | −41.0 points |
Both interventions cut false positives massively and by nearly the same amount. The preregistered primary comparison — spirals against matched ordinary galaxies — came out at +0.7 points in favor of spirals, and inconsistent in sign across architectures (it favored spirals for two, and reversed for the third). The apparent spiral effect was mostly a data-quantity effect: under this protocol, the look-alike morphology was not necessary.
Verification
The primary and secondary endpoints of the 90-run study were re-derived by an independently written verifier that shared no aggregation or interval code: 13 exact matches, 2 differences traceable to Monte-Carlo precision, and no substantive discrepancy. That audit covers this experiment, not every number in the project — a later recount in the follow-up phase did move one reported figure, and I've kept both facts on the record.
The follow-up, killed by its own rule
A proposed extension asked whether negatives could be chosen smartly — targeting regions of galaxy space that are common in the real sky but thin in the training set. Before running it, the success criterion was written down and frozen. It was not a single number: the new method had to beat both baselines by at least 5 percentage points of macro false-positive rate at equal verified-negative count — or match the better one on half the review budget — and win on paired seed direction in at least 4 of 5 seeds, without tripping either of two recall disqualifiers. It was frozen on 2026-07-22, before a single label existed.
Blind labeling
Five competing selection methods each picked galaxies from a 9,805-object pool. The 182 unique picks were pooled, stripped of any indication of which method chose them, assigned random IDs (OBJ-001 through OBJ-182), and hand-labeled blind — verified non-lens, possible lens, ambiguous, or unusable — with a written reason recorded for every single object. The method-to-object mapping was sealed under a hash before any label existed, the labels were committed a week later, and the same hash reappears in the unblinding record — so the ordering is checkable from the commit history rather than taken on trust. It is a self-seal, though: one person holding both files, with no third-party timestamping. Final tally: 131 verified negatives, 21 ambiguous, 20 unusable, and 10 possible lenses, which were quarantined from all further use.
The blind labels showed the smart method picked the cleanest galaxies of the five: 45 of its 50 picks came back confirmed non-lenses, a 90% yield, against 38 of 50 (76%) for flux-stratified random selection. That is a statement about label quality only. It was explicitly not treated as partial credit toward the efficacy rule, which the method still had to pass on its own terms.
The kill-switch fires
Then came the 100-run training comparison — 100 runs, no failures, no retries. Against flux-stratified random selection the smart method improved the false-positive rate by 0.00 points; against density-weighted risk selection, by 0.20. The rule demanded at least 5 against both, and seed direction favored the new method in 0 of 5 seeds against the first baseline and 1 of 5 against the second. Part of the reason is that the baseline detector was already near zero false positives on the new test sets, leaving almost no room to improve. The preregistered rule failed, and development stopped as promised. That stop, more than any single number, is the part of this project I would defend hardest: the criterion existed before the results did, and it was honored.
What this does and doesn't show
Supported
- Under this protocol, the tested brightness- and redshift-matched ordinary-negative control reduced false positives about as much as the spiral look-alike condition did.
- The apparent spiral-specific effect from the earlier, uncontrolled comparison was mostly a data-quantity effect.
- The follow-up's preregistered success criterion was not met, and development stopped accordingly.
Not claimed
- Not that any matched negatives work — only the specific matched control tested here was evaluated.
- Not that broader negative coverage is the proven mechanism; it is a leading interpretation, not an established one.
- Not that the smart selection method works. Its cleaner label yield says its picks were easier to adjudicate, not that it produced a better detector — on the measure that was preregistered, it did not.
- Not that any lenses were discovered. The 10 "possible lens" flags are one person's blind visual triage of 52-arcsecond cutouts, with no modeling, no spectroscopy, and no expert review.
- Not that the detector is deployment-ready or that the false-positive problem is solved.