OPENPROOF · SCREENING DESK
run osf-A09
maybescreening

What parameter-estimation bias does each glitch-subtraction method introduce in gravitational-wave parameter estimation, measured as the offset between injected and recovered parameters on public LIGO-Virgo-KAGRA data?

sources given with the question (8)
Non-exhaustive sources that may help:
- Gravitational Wave Open Science Center -- calibrated strain and data-quality products, CC-BY-4.0 — https://gwosc.org/
- O4 ended 18 Nov 2025 after ~2.5 years; O5 dates not announced; an Intermediate Run 1 is planned to start between late Oct and mid-Nov 2026 — https://observing.docs.ligo.org/plan/
- GWTC-4.0 (released 26 Aug 2025, 128 new significant candidates from O4a); GWTC-5.0 (May 2026, 161 new candidates, ~400 total) — https://arxiv.org/abs/2605.27225
- O4b open data paper: 'Open Data from LIGO, Virgo, and KAGRA through the Second Part of the Fourth Observing Run', arXiv:2605.27090 — https://arxiv.org/abs/2605.27090
- Gravity Spy; Zevin M. et al., Class Quantum Grav 34(6):064003 (2017); ~25% of the 161 new GWTC-5.0 candidates required glitch mitigation — https://doi.org/10.1088/1361-6382/aa5cea
- Gravity Spy multi-view fusion for O4, Class Quantum Grav (2025) — https://doi.org/10.1088/1361-6382/adf58b
- Open analysis stack: GWpy 4.0.1 (GPL-3.0), PyCBC 2.11.0 (GPL-3.0), bilby 2.8.2 (MIT) — https://gwosc.org/software/
gate Declined · budget_exhaustedno proof reached the judges4 calls$1.331166.7s

For the human reviewer

Why maybe

The run did not establish openness: no proof was judged, every literature-search rung failed, and the cited BayesWave evidence already contains at least a partial injection study whose configurations and coverage results were misreported.

Why the gate declined

solver budget spent

Literature search (engine-witnessed · openalex · 2026-09-02 · counts link to the live query)

Why no proof was judged

formalizer — rejected the solver's answer
The solution contains substantive BayesWave errors. First, Table 2 for Blip glitch 1 has 19/21—not 17/21—pre-subtraction 90%-interval successes: chirp mass 5/7, mass ratio 7/7, and effective spin 7/7; post-subtraction is 21/21. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) Second, the row incorrectly combines “BayesWave Joint / CBC+Glitch.” The reported injection study used the model-agnostic Joint configuration, while CBC+Glitch is a distinct configuration discussed as producing a less severe bias in the difficult overlap case; therefore the Joint results, including the failure below 0.75 overlap, cannot be attributed indiscriminately to CBC+Glitch. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) The solution should separate those configurations and correct the coverage count.

Calls, in firing order

1. solver 119.24s · $0.802 web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search web_search

system

You are one role in an automated verified-reasoning pipeline. Your reply is parsed by a program: return exactly what is asked, nothing else. You have no shell or filesystem. Use the web_search tool to verify facts, citations, and definitions rather than relying on recall. Use the code interpreter (Python; sympy/numpy/scipy available) to recompute rather than trusting arithmetic.

Solve the given problem. Show your reasoning.

Use web search for anything you're even remotely unsure about.

user

What parameter-estimation bias does each glitch-subtraction method introduce in gravitational-wave parameter estimation, measured as the offset between injected and recovered parameters on public LIGO-Virgo-KAGRA data?

Non-exhaustive sources that may help:
- Gravitational Wave Open Science Center -- calibrated strain and data-quality products, CC-BY-4.0 — https://gwosc.org/
- O4 ended 18 Nov 2025 after ~2.5 years; O5 dates not announced; an Intermediate Run 1 is planned to start between late Oct and mid-Nov 2026 — https://observing.docs.ligo.org/plan/
- GWTC-4.0 (released 26 Aug 2025, 128 new significant candidates from O4a); GWTC-5.0 (May 2026, 161 new candidates, ~400 total) — https://arxiv.org/abs/2605.27225
- O4b open data paper: 'Open Data from LIGO, Virgo, and KAGRA through the Second Part of the Fourth Observing Run', arXiv:2605.27090 — https://arxiv.org/abs/2605.27090
- Gravity Spy; Zevin M. et al., Class Quantum Grav 34(6):064003 (2017); ~25% of the 161 new GWTC-5.0 candidates required glitch mitigation — https://doi.org/10.1088/1361-6382/aa5cea
- Gravity Spy multi-view fusion for O4, Class Quantum Grav (2025) — https://doi.org/10.1088/1361-6382/adf58b
- Open analysis stack: GWpy 4.0.1 (GPL-3.0), PyCBC 2.11.0 (GPL-3.0), bilby 2.8.2 (MIT) — https://gwosc.org/software/

reasoning

{
 "summary": [
  "**Researching data comparison methods**\n\nI think the user is asking for an empirical research comparison and whether I need to clarify the methods for that. They seek future data projected to August 31, 2026, so I need to do some web research. I should look into glitch subtraction methods like BayesWave, gwsubtract, and perhaps gating or inpainting. It seems they want to analyze parameter-estimation bias on public LVK data, but I need specifics to proceed. I'll identify relevant benchmark literature and search the web for details."
 ]
}

web_search

{
 "action": {
  "type": "search",
  "queries": [
   "gravitational wave glitch subtraction parameter estimation injected recovered offset BayesWave gwsubtract public data",
   "site:arxiv.org glitch mitigation parameter estimation injections BayesWave gwsubtract gravitational wave",
   "LIGO glitch subtraction methods parameter estimation bias injection recovered parameters",
   "public LIGO data glitch subtraction benchmark parameter estimation bias"
  ],
  "query": "gravitational wave glitch subtraction parameter estimation injected recovered offset BayesWave gwsubtract public data"
 }
}

reasoning

{
 "summary": [
  "**Exploring parameter estimation bias**\n\nI see the question asks about the parameter-estimation bias from glitch-subtraction methods, particularly how it relates to offsets between injected and recovered parameters in public data. It seems like there's an expectation for actual numbers, but there aren't clear experiment definitions. I should explain what \"underdetermined\" means, and I notice there’s no universal method bias outlined in the literature. Summarizing known numerical offsets from injection papers could be beneficial, so I’ll need to inspect a relevant paper."
 ]
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/abs/2311.09159"
 }
}

reasoning

{
 "summary": [
  "**Examining public data methods**\n\nI need to gather some numeric figures that aren't currently represented. I could try fetching PDFs or source text, but they might not be available through open access. The user's request seems to suggest a research task, expecting a table of biases for methods applied to public data. It's important to clarify that there’s no universal number for this and establish a benchmark. I should also look at the reported accuracy values from the literature and inspect figures 12 and 13 for relevant data."
 ]
}

web_search

{
 "action": {
  "pattern": "Figure 12",
  "type": "find_in_page",
  "url": "https://arxiv.org/abs/2505.00657"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/abs/2505.00657"
 }
}

reasoning

{
 "summary": [
  "**Considering options for plots**\n\nI'm thinking about how I could click on the PDF and take screenshots of the pages that have the plots. It seems like a practical solution, but I want to be sure I capture everything clearly. I wonder if there’s a better way to extract the plots without losing quality, making it easier to use them later. Hmm, there’s a lot to consider here! But for now, screenshots seem to be a straightforward approach."
 ]
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/abs/2505.00657"
 }
}

reasoning

{
 "summary": [
  "**Considering PDF Analysis**\n\nI’m thinking about inspecting relevant lines in PDFs. I need to focus on extracting the information that’s important. I wonder how I can methodically check for details to ensure I get the right context. There might be tools or methods that can help streamline this process. It seems crucial to stay organized while working through the content, especially if there’s a lot to sift through. Let's figure out the best approach!"
 ]
}

web_search

{
 "action": {
  "pattern": "standard accuracy",
  "type": "find_in_page",
  "url": "https://arxiv.org/pdf/2505.00657"
 }
}

reasoning

{
 "summary": [
  "**Extracting visual data**\n\nI'm considering taking a screenshot of page 11 to visually extract the bars. This sounds like a practical approach! I want to ensure that I capture the detail well, so I'll focus on getting a clear image that can show what I need accurately. After that, I’ll analyze the bars for the information I’m looking for. It's really about making sure I represent the data effectively, so this sounds like a workable plan."
 ]
}

web_search

{
 "action": null
}

web_search

{
 "action": null
}

reasoning

{
 "summary": [
  "**Investigating glitch-subtraction methods**\n\nI’m thinking about how we can read approximate bars. The user mentioned “each glitch-subtraction method,” which might include techniques like BayesWave and gwsubtract, especially related to LVK. However, it seems there might not be documentation on injection-recovery offsets for gwsubtract in the literature. I should consider doing a more targeted search to verify this and gather more information for clarity!"
 ]
}

web_search

{
 "action": {
  "type": "search",
  "queries": [
   "gwsubtract injection parameter estimation bias recovered injected parameters",
   "\"gwsubtract\" \"parameter estimation\" injection",
   "\"BayesWave\" \"gwsubtract\" bias injection recovery",
   "auxiliary channel glitch subtraction injection parameter estimation bias LIGO"
  ],
  "query": "gwsubtract injection parameter estimation bias recovered injected parameters"
 }
}

reasoning

{
 "summary": [
  "**Analyzing glitch-subtraction methods**\n\nI'm thinking about the various glitch-subtraction methods mentioned in literature, like BayesWave, gwsubtract, SHAPES, and others. The question might be asking for all these methods benchmarked on public data. However, there's no apples-to-apples comparison for each method since only some have specific studies with different setups and metrics. gwsubtract can't really be applied to standard public strain data either due to limitations. I should look into the BayesWave paper for the numeric metrics provided."
 ]
}

web_search

{
 "action": {
  "pattern": "bias statistic",
  "type": "find_in_page",
  "url": "https://arxiv.org/pdf/2311.09159"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/pdf/2311.09159"
 }
}

web_search

{
 "action": {
  "type": "search",
  "queries": [
   "\"Assessing and mitigating\" glitches \"credible interval\" BayesWave bias",
   "site:journals.aps.org/prd/abstract/10.1103/PhysRevD.110.122002",
   "\"PhysRevD.110.122002\" parameter bias glitch"
  ],
  "query": "\"Assessing and mitigating\" glitches \"credible interval\" BayesWave bias"
 }
}

reasoning

{
 "summary": [
  "**Considering figure inspection**\n\nI need to take a closer look at the BayesWave figure pages, possibly taking screenshots of all 21 pages. The relevant figures likely start from page 11 onward. I should find the captions using OCR, but since there's a timeout issue, I might have to open the web parsed pages without text. I'll take screenshots from pages 10 to 13 one by one to make sure I'm capturing the necessary information."
 ]
}

web_search

{
 "action": null
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://authors.library.caltech.edu/records/pb1vg-rf514"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://authors.library.caltech.edu/records/pb1vg-rf514"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://authors.library.caltech.edu/records/pb1vg-rf514"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/abs/2311.09159"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/html/2311.09159v2"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/html/2311.09159v2"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/abs/1806.00532"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/html/2311.09159v2"
 }
}

reasoning

{
 "summary": [
  "**Evaluating glitch-subtraction methods**\n\nI'm considering the question about \"each glitch-subtraction method,\" which may relate only to LVK operational methods. I think mentioning a \"No subtraction\" comparator could be relevant. It seems the impact of certain shapes is negligible, but it’s not really an apples-to-apples comparison. I should explain that gating or frequency cuts are more about mitigation than actual subtraction. If the user wants a benchmark, a new campaign might be necessary, and I need to provide citations for any factual claims."
 ]
}

output

There is **no published, apples-to-apples number for “the bias introduced by each glitch-subtraction method” on public LVK data**. The offset is not a property of the subtraction algorithm alone: it depends on glitch class and SNR, signal morphology and SNR, signal–glitch time separation, detector network, waveform model, priors, and the particular glitch reconstruction drawn from its posterior.

Define, for parameter \(\theta\),

\[
\Delta\theta=\hat{\theta}-\theta_{\rm inj},
\qquad
z_\theta=\frac{|\hat{\theta}-\theta_{\rm inj}|}{\sigma_\theta},
\]

where \(\hat{\theta}\) must be specified—for example, posterior median or maximum-likelihood value. A method’s systematic bias should be estimated from an ensemble, such as \(\langle\Delta\theta\rangle\), rather than from one injection.

## What has actually been measured

| Method | Public-data injection result | Defensible statement about subtraction-induced bias |
|---|---|---|
| **BayesWave Joint / CBC+Glitch** | Software BBH injections of network SNR 15 were placed near real O2/O3 Blip, Tomte, and scattered-light glitches. After successful subtraction, injected chirp mass, mass ratio, and \(\chi_{\rm eff}\) were generally inside the recovered 90% intervals, with injected–recovered waveform overlap \(>0.9\). Blip contamination matters mainly within roughly \(0.05\)–\(0.1\) s of merger; Tomtes can produce severe pre-subtraction bias; the tested scattering glitches produced little bias even before subtraction. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) | **Usually consistent with zero at posterior resolution**, but not guaranteed. One Blip case at a \(0.025\)-s separation remained biased after subtraction: chirp mass and \(\chi_{\rm eff}\) excluded the injection, and waveform overlap fell below \(0.75\). Thus BayesWave can occasionally leave residual-glitch bias or misallocate power between signal and glitch. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) |
| **gwsubtract / witness-channel linear subtraction** | This method estimates a transfer function from an auxiliary witness to strain. The published O3 study gives event case studies and glitch-model uncertainties, not a controlled table of injected-minus-recovered CBC parameters directly comparable with BayesWave. ([arxiv.org](https://arxiv.org/abs/2207.03429)) | **No method-specific numerical PE bias has been established on a common public glitch-injection benchmark.** It should be close to zero when the witness is signal-safe and the transfer function is correct, but under-subtraction and transfer-function error remain possible. The O3 auxiliary-channel release makes an independent test possible. ([gwosc.org](https://gwosc.org/O3/auxiliary/?utm_source=openai)) |
| **Normalizing-flow, glitch-class-informed joint model** | A Blip-trained model was tested using signals injected into real O1 and O3 Gravity Spy glitches. The study used “standard accuracy” \(z_\theta=|\theta_{\max L}-\theta_{\rm inj}|/\sigma_\theta\). After subtraction, mean \(z_\theta\) was approximately \(0.4\)–\(2.0\) for O1 Blips and \(0.9\)–\(1.2\) for O3 Blips, depending on parameter; contaminated-data values reached about \(17\) and \(12\), respectively. ([arxiv.org](https://arxiv.org/pdf/2505.00657)) | **Residual offsets are approximately one posterior standard deviation on average**, hence broadly statistically consistent with unbiased recovery. The largest remaining O1 value was about \(2\sigma\), for right ascension. These are absolute normalized offsets, not signed biases. |
| **Blip-trained normalizing flow applied to Tomtes** | For 20 O3 Tomtes, approximate post-subtraction \(z_\theta\) values were: \(q\sim1.3\), \(\mathcal M\sim1.0\), \(\theta_{JN}\sim1.0\), \(d_L\sim1.1\), RA \(\sim1.0\), and Dec \(\sim0.8\). Before subtraction they were approximately \(1.6,1.8,5.5,4.0,1.9,1.2\). ([arxiv.org](https://arxiv.org/pdf/2505.00657)) | **About \(1\sigma\) residual offset on average**, although model selection correctly recognized the mismatched Blip model poorly: only 3 of 20 signal-plus-Tomte cases preferred it. Thus reduced PE offset does not establish that the glitch model is correct. ([arxiv.org](https://arxiv.org/abs/2505.00657)) |
| **SHAPES adaptive-spline subtraction** | Tested on GW170817 and GW170817-like BNS injections overlapping broadband glitches; the reported effect of subtraction on the GW signal was negligible. ([arxiv.org](https://arxiv.org/abs/2301.02398?utm_source=openai)) | **Consistent with negligible induced bias in that limited BNS test**, but no population-level \(\langle\Delta\theta\rangle\) comparable to the numbers above was published. |
| **No subtraction** | Real broadband Blip and Tomte glitches can move posteriors many standard deviations from the injection; the normalizing-flow tests reached mean \(z_\theta\sim12\)–17 for some extrinsic parameters. Narrowband/scattering glitches with sufficiently different morphology produced little change in the tested BBHs. ([arxiv.org](https://arxiv.org/pdf/2505.00657)) | This measures **glitch-induced**, not subtraction-induced, bias. It is highly morphology- and timing-dependent. |

## BayesWave coverage numbers

For the published BayesWave study’s first Blip, all \(21/21\) combinations of seven injection times and three intrinsic parameters contained the true value after subtraction, versus \(17/21\) before subtraction. For the second Blip, \(18/21\) contained the truth after subtraction, versus \(9/21\) before. The failures were concentrated in the particularly difficult \(0.025\)-s overlap. ([arxiv.org](https://arxiv.org/html/2311.09159v2))

For the illustrated Tomte, all \(27/27\) post-subtraction intervals contained the injection, compared with only \(1/27\) before subtraction. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) These are **coverage counts**, not mean signed offsets.

## Bottom line

- **BayesWave:** normally leaves no statistically significant offset, but rare near-coincident failures can remain strongly biased.
- **gwsubtract:** no public, controlled injection study currently provides a directly comparable parameter-offset vector; its numerical bias is therefore **not established**, rather than demonstrably zero.
- **Class-informed normalizing-flow subtraction:** approximately \(1\sigma\) mean absolute residual offset on O3 tests, with somewhat larger values—up to about \(2\sigma\)—on O1.
- **SHAPES:** negligible measured impact in its limited GW170817-like tests, without a broad numerical bias survey.
- Consequently, no evidence supports assigning a single fixed offset such as \(\Delta\mathcal M\), \(\Delta q\), or \(\Delta\chi_{\rm eff}\) to any method.

A valid method comparison would require the same GPS segments, injection population, waveform and priors, sampler, posterior summary, glitch reconstruction convention, and random seeds for every method, followed by reporting \(\langle\Delta\theta\rangle\), RMS error, normalized bias, and empirical credible-interval coverage. Existing publications do not yet provide that common benchmark.
2. formalizer 47.42s · $0.529 web_search web_search web_search web_search web_search web_search web_search web_search web_search

system

You are one role in an automated verified-reasoning pipeline. Your reply is parsed by a program: return exactly what is asked, nothing else. You have no shell or filesystem. Use the web_search tool to verify facts, citations, and definitions rather than relying on recall.

You formalize solutions into proofs. The solutions you receive were
generated by an LLM and may contain hallucinations, wrong facts, or
flawed reasoning: treat every claim as potentially wrong, and do not
assume the solver's reasoning is valid unless you can justify it.

On every call, you must choose one of two actions:

- action="proof": Produce a formal proof for the solution (see
  proof format below). This is the default — start by trying to
  formalize.
- action="reject": If while formalizing you find any substantive
  issues with the solution (incorrect facts, flawed reasoning,
  missing cases, etc.), reject the solution and set reject_reason
  to explain what's wrong; the solver will get this and retry.

A proof is a sequence of states. Each state is a list of strings.
Index 0 is the expression being solved for. It starts as a goal
variable (like "ANSWER") and ends when resolved.

Each step has: state (full list of strings after this step),
justification_type ("citation", "problem_given", or "computation"),
and justification (why).

One transformation per step. Every step must be justified.

When a step's justification is a named mode of inference, use
justification_type: citation. The cited "result" is the mode of
inference itself.

PREMISES MUST BE EXPLICIT. When a step relies on a premise that
isn't already in the previous state — including bounds, conditions,
edge cases, factual claims, or assumptions like "X is a Y" — you
MUST first add a step that introduces the premise explicitly. That
added step needs its own justification. Do not bury premises in
justification text. If you cannot justify a premise, your proof
must handle the case without it (e.g., by considering all cases).

Use web search for anything you're even remotely unsure about.

user

Problem: What parameter-estimation bias does each glitch-subtraction method introduce in gravitational-wave parameter estimation, measured as the offset between injected and recovered parameters on public LIGO-Virgo-KAGRA data?

Non-exhaustive sources that may help:
- Gravitational Wave Open Science Center -- calibrated strain and data-quality products, CC-BY-4.0 — https://gwosc.org/
- O4 ended 18 Nov 2025 after ~2.5 years; O5 dates not announced; an Intermediate Run 1 is planned to start between late Oct and mid-Nov 2026 — https://observing.docs.ligo.org/plan/
- GWTC-4.0 (released 26 Aug 2025, 128 new significant candidates from O4a); GWTC-5.0 (May 2026, 161 new candidates, ~400 total) — https://arxiv.org/abs/2605.27225
- O4b open data paper: 'Open Data from LIGO, Virgo, and KAGRA through the Second Part of the Fourth Observing Run', arXiv:2605.27090 — https://arxiv.org/abs/2605.27090
- Gravity Spy; Zevin M. et al., Class Quantum Grav 34(6):064003 (2017); ~25% of the 161 new GWTC-5.0 candidates required glitch mitigation — https://doi.org/10.1088/1361-6382/aa5cea
- Gravity Spy multi-view fusion for O4, Class Quantum Grav (2025) — https://doi.org/10.1088/1361-6382/adf58b
- Open analysis stack: GWpy 4.0.1 (GPL-3.0), PyCBC 2.11.0 (GPL-3.0), bilby 2.8.2 (MIT) — https://gwosc.org/software/


Solution: There is **no published, apples-to-apples number for “the bias introduced by each glitch-subtraction method” on public LVK data**. The offset is not a property of the subtraction algorithm alone: it depends on glitch class and SNR, signal morphology and SNR, signal–glitch time separation, detector network, waveform model, priors, and the particular glitch reconstruction drawn from its posterior.

Define, for parameter \(\theta\),

\[
\Delta\theta=\hat{\theta}-\theta_{\rm inj},
\qquad
z_\theta=\frac{|\hat{\theta}-\theta_{\rm inj}|}{\sigma_\theta},
\]

where \(\hat{\theta}\) must be specified—for example, posterior median or maximum-likelihood value. A method’s systematic bias should be estimated from an ensemble, such as \(\langle\Delta\theta\rangle\), rather than from one injection.

## What has actually been measured

| Method | Public-data injection result | Defensible statement about subtraction-induced bias |
|---|---|---|
| **BayesWave Joint / CBC+Glitch** | Software BBH injections of network SNR 15 were placed near real O2/O3 Blip, Tomte, and scattered-light glitches. After successful subtraction, injected chirp mass, mass ratio, and \(\chi_{\rm eff}\) were generally inside the recovered 90% intervals, with injected–recovered waveform overlap \(>0.9\). Blip contamination matters mainly within roughly \(0.05\)–\(0.1\) s of merger; Tomtes can produce severe pre-subtraction bias; the tested scattering glitches produced little bias even before subtraction. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) | **Usually consistent with zero at posterior resolution**, but not guaranteed. One Blip case at a \(0.025\)-s separation remained biased after subtraction: chirp mass and \(\chi_{\rm eff}\) excluded the injection, and waveform overlap fell below \(0.75\). Thus BayesWave can occasionally leave residual-glitch bias or misallocate power between signal and glitch. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) |
| **gwsubtract / witness-channel linear subtraction** | This method estimates a transfer function from an auxiliary witness to strain. The published O3 study gives event case studies and glitch-model uncertainties, not a controlled table of injected-minus-recovered CBC parameters directly comparable with BayesWave. ([arxiv.org](https://arxiv.org/abs/2207.03429)) | **No method-specific numerical PE bias has been established on a common public glitch-injection benchmark.** It should be close to zero when the witness is signal-safe and the transfer function is correct, but under-subtraction and transfer-function error remain possible. The O3 auxiliary-channel release makes an independent test possible. ([gwosc.org](https://gwosc.org/O3/auxiliary/?utm_source=openai)) |
| **Normalizing-flow, glitch-class-informed joint model** | A Blip-trained model was tested using signals injected into real O1 and O3 Gravity Spy glitches. The study used “standard accuracy” \(z_\theta=|\theta_{\max L}-\theta_{\rm inj}|/\sigma_\theta\). After subtraction, mean \(z_\theta\) was approximately \(0.4\)–\(2.0\) for O1 Blips and \(0.9\)–\(1.2\) for O3 Blips, depending on parameter; contaminated-data values reached about \(17\) and \(12\), respectively. ([arxiv.org](https://arxiv.org/pdf/2505.00657)) | **Residual offsets are approximately one posterior standard deviation on average**, hence broadly statistically consistent with unbiased recovery. The largest remaining O1 value was about \(2\sigma\), for right ascension. These are absolute normalized offsets, not signed biases. |
| **Blip-trained normalizing flow applied to Tomtes** | For 20 O3 Tomtes, approximate post-subtraction \(z_\theta\) values were: \(q\sim1.3\), \(\mathcal M\sim1.0\), \(\theta_{JN}\sim1.0\), \(d_L\sim1.1\), RA \(\sim1.0\), and Dec \(\sim0.8\). Before subtraction they were approximately \(1.6,1.8,5.5,4.0,1.9,1.2\). ([arxiv.org](https://arxiv.org/pdf/2505.00657)) | **About \(1\sigma\) residual offset on average**, although model selection correctly recognized the mismatched Blip model poorly: only 3 of 20 signal-plus-Tomte cases preferred it. Thus reduced PE offset does not establish that the glitch model is correct. ([arxiv.org](https://arxiv.org/abs/2505.00657)) |
| **SHAPES adaptive-spline subtraction** | Tested on GW170817 and GW170817-like BNS injections overlapping broadband glitches; the reported effect of subtraction on the GW signal was negligible. ([arxiv.org](https://arxiv.org/abs/2301.02398?utm_source=openai)) | **Consistent with negligible induced bias in that limited BNS test**, but no population-level \(\langle\Delta\theta\rangle\) comparable to the numbers above was published. |
| **No subtraction** | Real broadband Blip and Tomte glitches can move posteriors many standard deviations from the injection; the normalizing-flow tests reached mean \(z_\theta\sim12\)–17 for some extrinsic parameters. Narrowband/scattering glitches with sufficiently different morphology produced little change in the tested BBHs. ([arxiv.org](https://arxiv.org/pdf/2505.00657)) | This measures **glitch-induced**, not subtraction-induced, bias. It is highly morphology- and timing-dependent. |

## BayesWave coverage numbers

For the published BayesWave study’s first Blip, all \(21/21\) combinations of seven injection times and three intrinsic parameters contained the true value after subtraction, versus \(17/21\) before subtraction. For the second Blip, \(18/21\) contained the truth after subtraction, versus \(9/21\) before. The failures were concentrated in the particularly difficult \(0.025\)-s overlap. ([arxiv.org](https://arxiv.org/html/2311.09159v2))

For the illustrated Tomte, all \(27/27\) post-subtraction intervals contained the injection, compared with only \(1/27\) before subtraction. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) These are **coverage counts**, not mean signed offsets.

## Bottom line

- **BayesWave:** normally leaves no statistically significant offset, but rare near-coincident failures can remain strongly biased.
- **gwsubtract:** no public, controlled injection study currently provides a directly comparable parameter-offset vector; its numerical bias is therefore **not established**, rather than demonstrably zero.
- **Class-informed normalizing-flow subtraction:** approximately \(1\sigma\) mean absolute residual offset on O3 tests, with somewhat larger values—up to about \(2\sigma\)—on O1.
- **SHAPES:** negligible measured impact in its limited GW170817-like tests, without a broad numerical bias survey.
- Consequently, no evidence supports assigning a single fixed offset such as \(\Delta\mathcal M\), \(\Delta q\), or \(\Delta\chi_{\rm eff}\) to any method.

A valid method comparison would require the same GPS segments, injection population, waveform and priors, sampler, posterior summary, glitch reconstruction convention, and random seeds for every method, followed by reporting \(\langle\Delta\theta\rangle\), RMS error, normalized bias, and empirical credible-interval coverage. Existing publications do not yet provide that common benchmark.

Formalize this into a proof, or reject if it has errors.

reasoning

{
 "summary": [
  "**Verifying sources**\n\nI need to verify everything thoroughly. If there are any errors, I’ll likely need to reject that information. I should inspect the cited papers closely, making sure they are primary sources. I think I need to search the web to double-check what I find and ensure accuracy. It’s important to be meticulous about these details to provide the user with reliable and trustworthy information! A careful approach will definitely help!"
 ]
}

web_search

{
 "action": {
  "type": "search",
  "queries": [
   "site:arxiv.org 2311.09159 glitch subtraction BayesWave parameter estimation blip tomte scattered light",
   "site:arxiv.org 2505.00657 normalizing flow glitch subtraction gravitational wave parameter estimation",
   "site:arxiv.org 2301.02398 SHAPES glitch subtraction GW170817",
   "site:arxiv.org 2207.03429 gwsubtract O3 auxiliary witness channel"
  ],
  "query": "site:arxiv.org 2311.09159 glitch subtraction BayesWave parameter estimation blip tomte scattered light"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/html/2311.09159v2"
 }
}

web_search

{
 "action": {
  "pattern": "21/21",
  "type": "find_in_page",
  "url": "https://arxiv.org/html/2311.09159v2"
 }
}

reasoning

{
 "summary": [
  "**Checking counts and figures**\n\nI'm noticing a discrepancy in the counts from tables. For Blip1, it should show 7 counts instead of 9. I counted 19 before and 21 after, while the solver claims 17, which seems incorrect. Then, looking at Table 3, I’m confirming that counts are also off with 4 before and 9 after chirp, leading me to think there might be a typo or inconsistency. I need to inspect the figures further to clarify the strong factual error."
 ]
}

web_search

{
 "action": {
  "pattern": "standard accuracy",
  "type": "find_in_page",
  "url": "https://arxiv.org/pdf/2505.00657"
 }
}

web_search

{
 "action": {
  "type": "open_page",
  "url": "https://arxiv.org/pdf/2505.00657"
 }
}

reasoning

{
 "summary": [
  "**Evaluating figure accuracy**\n\nI need to get a screenshot of the approximate bars from the figures. I want to trust the values, but there's a chance I could be wrong. Maybe I should focus on screenshotting page 11, which I think contains figures 12 and 13? Since the PDF pages count from zero, I’m ensuring I'm referring to the correct index. It’s a little tricky, but I’m working through it!"
 ]
}

web_search

{
 "action": null
}

web_search

{
 "action": null
}

web_search

{
 "action": {
  "pattern": "injection",
  "type": "find_in_page",
  "url": "https://arxiv.org/abs/2301.02398"
 }
}

reasoning

{
 "summary": [
  "**Analyzing method discrepancies**\n\nI’m noticing a potential issue with \"SHAPES adaptive spline subtraction,\" and I wonder if \"SHAPES\" is indeed an acronym. I need to investigate if this could solely reject counts. \n\nThere appears to be a conceptual issue regarding claims about residual offsets. The user wants to know about methods, and while they're defined, z values are provided instead. \n\nI also see that \"BayesWave Joint / CBC+Glitch\" conflates two methods. Study 2311 only uses Joint, despite the source mentioning CBC+Glitch as more effective. This could be misleading in representation.",
  "**Clarifying method distinctions**\n\nI’m noticing that the bottom clearly states BayesWave without a separate mention of CBC+Glitch. Since there's a question about \"each glitch subtraction method,\" it's important to distinguish between them. The indicated \"standard LVK deglitching\" draws from random posteriors. \n\nIt seems like the solution mentioning BayesWave Joint / CBC+Glitch may not accurately reflect distinctions. I should check for the latest papers to see if there's a benchmark after May 2025, specifically regarding the current claims. I need to search for updates from 2025 and 2026."
 ]
}

web_search

{
 "action": {
  "type": "search",
  "queries": [
   "2026 gravitational wave glitch subtraction methods parameter estimation injection comparison BayesWave gwsubtract SHAPES normalizing flow",
   "2025 gravitational wave glitch mitigation parameter estimation benchmark public data subtraction comparison",
   "site:arxiv.org gravitational wave glitch subtraction parameter estimation injected recovered 2026",
   "site:arxiv.org gravitational wave deglitching methods comparison parameter estimation public data"
  ],
  "query": "2026 gravitational wave glitch subtraction methods parameter estimation injection comparison BayesWave gwsubtract SHAPES normalizing flow"
 }
}

output

{
 "action": "reject",
 "proof": {
  "initial_state": [],
  "steps": []
 },
 "reject_reason": "The solution contains substantive BayesWave errors. First, Table 2 for Blip glitch 1 has 19/21—not 17/21—pre-subtraction 90%-interval successes: chirp mass 5/7, mass ratio 7/7, and effective spin 7/7; post-subtraction is 21/21. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) Second, the row incorrectly combines “BayesWave Joint / CBC+Glitch.” The reported injection study used the model-agnostic Joint configuration, while CBC+Glitch is a distinct configuration discussed as producing a less severe bias in the difficult overlap case; therefore the Joint results, including the failure below 0.75 overlap, cannot be attributed indiscriminately to CBC+Glitch. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) The solution should separate those configurations and correct the coverage count."
}
3. searcher 21.22s · $0.049

system

You are one role in an automated verified-reasoning pipeline. Your reply is parsed by a program: return exactly what is asked, nothing else. You have no shell or filesystem.

You write the literature-search queries for one verification run, from its
digest: the question (a research proposal's hypothesis), the outcome, the
proof, and every audited claim with the checker's reason. You do not search.
An engine runs every query you return against OpenAlex (title, abstract and
full text of ~250M works) and records the exact URL, the total match count
and the top hits; the adjudicator reads those results, not your opinion.

Return one `searches` entry per claim worth testing:
- every absence claim the run relied on or failed on ("no study reports X",
  "X has not been measured", "no benchmark covers Y"), in the digest's words;
- the proposal's question itself: has it already been answered?

For each, `queries` is a ladder from broad to narrow, as many rungs as the
vocabulary needs. The broadest rung names the topic and must plausibly match
published work; each narrower rung adds a condition the claim turns on.
Cover the field's synonyms and named artefacts (tools, datasets, benchmarks,
versions) across rungs rather than inside one query. A ladder whose broad
rung matches nothing is uncalibrated and proves nothing, so prefer several
plain rungs to one long one.

Query syntax: whole words, stemmed; AND / OR / NOT (uppercase) and
parentheses; "double quotes" for phrases. `from_date` / `to_date` bound the
publication date (YYYY-MM-DD; "" for unbounded). `issns`: journal ISSNs to
restrict to, only when the claim names venues; otherwise [].

user

QUESTION:
What parameter-estimation bias does each glitch-subtraction method introduce in gravitational-wave parameter estimation, measured as the offset between injected and recovered parameters on public LIGO-Virgo-KAGRA data?

Non-exhaustive sources that may help:
- Gravitational Wave Open Science Center -- calibrated strain and data-quality products, CC-BY-4.0 — https://gwosc.org/
- O4 ended 18 Nov 2025 after ~2.5 years; O5 dates not announced; an Intermediate Run 1 is planned to start between late Oct and mid-Nov 2026 — https://observing.docs.ligo.org/plan/
- GWTC-4.0 (released 26 Aug 2025, 128 new significant candidates from O4a); GWTC-5.0 (May 2026, 161 new candidates, ~400 total) — https://arxiv.org/abs/2605.27225
- O4b open data paper: 'Open Data from LIGO, Virgo, and KAGRA through the Second Part of the Fourth Observing Run', arXiv:2605.27090 — https://arxiv.org/abs/2605.27090
- Gravity Spy; Zevin M. et al., Class Quantum Grav 34(6):064003 (2017); ~25% of the 161 new GWTC-5.0 candidates required glitch mitigation — https://doi.org/10.1088/1361-6382/aa5cea
- Gravity Spy multi-view fusion for O4, Class Quantum Grav (2025) — https://doi.org/10.1088/1361-6382/adf58b
- Open analysis stack: GWpy 4.0.1 (GPL-3.0), PyCBC 2.11.0 (GPL-3.0), bilby 2.8.2 (MIT) — https://gwosc.org/software/

OUTCOME: Declined budget_exhausted
DETAIL: solver budget spent
NO PROOF WAS JUDGED — the formalizer rejected the solver's answer:
The solution contains substantive BayesWave errors. First, Table 2 for Blip glitch 1 has 19/21—not 17/21—pre-subtraction 90%-interval successes: chirp mass 5/7, mass ratio 7/7, and effective spin 7/7; post-subtraction is 21/21. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) Second, the row incorrectly combines “BayesWave Joint / CBC+Glitch.” The reported injection study used the model-agnostic Joint configuration, while CBC+Glitch is a distinct configuration discussed as producing a less severe bias in the difficult overlap case; therefore the Joint results, including the failure below 0.75 overlap, cannot be attributed indiscriminately to CBC+Glitch. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) The solution should separate those configurations and correct the coverage count.
TOOL ACTIVITY: formalizer:web_search x9, solver:web_search x21
STATS: 3 calls, 166.7s

reasoning

{
 "summary": [
  "**Clarifying claim requirements**\n\nI’m looking at the wording on audited claims that requires a checker's reason, but the inclusion rules are saying we need absence claims plus a proposal question, meaning we shouldn't count positives separately. I could frame the claim more as, \"Whether published work has already quantified...\" without asserting there’s no study, because that's actually the proposal question itself. It sounds good, and I might want to include some condition about \"public LIGO-Virgo-KAGRA data\" as the final piece."
 ]
}

output

{
 "searches": [
  {
   "claim": "What parameter-estimation bias does each glitch-subtraction method introduce in gravitational-wave parameter estimation, measured as the offset between injected and recovered parameters on public LIGO-Virgo-KAGRA data?",
   "queries": [
    "\"gravitational wave\" AND glitch AND \"parameter estimation\"",
    "\"gravitational wave\" AND (\"glitch subtraction\" OR \"glitch mitigation\") AND \"parameter estimation\"",
    "\"gravitational wave\" AND (\"glitch subtraction\" OR \"glitch mitigation\") AND (injection OR injected) AND (recovery OR recovered)",
    "\"gravitational wave\" AND glitch AND \"parameter bias\" AND injection",
    "\"gravitational wave\" AND glitch AND \"parameter estimation\" AND (coverage OR \"credible interval\")",
    "LIGO AND glitch AND \"parameter estimation\" AND \"public data\"",
    "GWOSC AND glitch AND \"parameter estimation\" AND injection",
    "BayesWave AND glitch AND \"parameter estimation\" AND injection",
    "BayesWave AND Joint AND glitch AND \"parameter estimation\"",
    "BayesWave AND \"CBC+Glitch\" AND \"parameter estimation\"",
    "gwsubtract AND glitch AND \"parameter estimation\"",
    "(\"gating\" OR inpainting) AND glitch AND \"gravitational wave\" AND \"parameter estimation\"",
    "(\"glitch subtraction\" OR \"glitch mitigation\") AND \"injected and recovered\" AND (LIGO OR Virgo OR KAGRA)",
    "(\"glitch subtraction\" OR \"glitch mitigation\") AND (bias OR offset) AND injection AND recovered AND (GWOSC OR \"public data\")"
   ],
   "from_date": "",
   "to_date": "",
   "issns": []
  }
 ]
}
4. adjudicator 5.65s · $0.054

system

You are one role in an automated verified-reasoning pipeline. Your reply is parsed by a program: return exactly what is asked, nothing else. You have no shell or filesystem. Use the web_search tool to verify facts, citations, and definitions rather than relying on recall. Use the code interpreter (Python; sympy/numpy/scipy available) to recompute rather than trusting arithmetic.

You give the screening verdict on one research-proposal question, using the
digest of a verification run: the question (the proposal's hypothesis), the
outcome, the proof, and every audited claim with the checker's full reason.

`verdict` — is the proposal's question genuinely open and worth a human
reviewer's time?
- yes: the run establishes that the question is real, open, and well-posed —
  a certified "not settled by the reviewed evidence", or a checker-verified
  absence of the result the proposal would supply.
- no: the run shows it is not a fundable open question — already settled by
  the literature, ill-posed, or its supporting claims collapse on checkable
  facts.
- maybe: the run leaves specific uncertainties only a human can resolve. If
  openness rests on something the run did not check — whether the analysis
  is already published, whether the data exists — that is maybe, with the
  check as a review item, not yes.

The digest may end with a LITERATURE SEARCH section: engine-run OpenAlex
queries with their total match counts, per claim, broad to narrow; a rung
with at most 10 matches lists them (title, year, venue, doi). Read it as evidence, not as a verdict. Hits whose titles answer the
proposal's question support no (already settled). Zero hits on the narrow
rungs of a calibrated ladder (its broad rungs matched) support yes for that
claim's absence. An uncalibrated ladder establishes nothing, and a FAILED
rung is unknown, not zero. Name the query or hit you rely on.

`explanation`: for yes or no, 2-4 sentences grounded only in the digest.
For maybe, one sentence naming the core uncertainty.

`review_items`: for maybe only — 2 to 6 concrete questions or checks for the
human reviewer, each answerable and each tied to something in the digest.
Empty for yes and no.

user

QUESTION:
What parameter-estimation bias does each glitch-subtraction method introduce in gravitational-wave parameter estimation, measured as the offset between injected and recovered parameters on public LIGO-Virgo-KAGRA data?

Non-exhaustive sources that may help:
- Gravitational Wave Open Science Center -- calibrated strain and data-quality products, CC-BY-4.0 — https://gwosc.org/
- O4 ended 18 Nov 2025 after ~2.5 years; O5 dates not announced; an Intermediate Run 1 is planned to start between late Oct and mid-Nov 2026 — https://observing.docs.ligo.org/plan/
- GWTC-4.0 (released 26 Aug 2025, 128 new significant candidates from O4a); GWTC-5.0 (May 2026, 161 new candidates, ~400 total) — https://arxiv.org/abs/2605.27225
- O4b open data paper: 'Open Data from LIGO, Virgo, and KAGRA through the Second Part of the Fourth Observing Run', arXiv:2605.27090 — https://arxiv.org/abs/2605.27090
- Gravity Spy; Zevin M. et al., Class Quantum Grav 34(6):064003 (2017); ~25% of the 161 new GWTC-5.0 candidates required glitch mitigation — https://doi.org/10.1088/1361-6382/aa5cea
- Gravity Spy multi-view fusion for O4, Class Quantum Grav (2025) — https://doi.org/10.1088/1361-6382/adf58b
- Open analysis stack: GWpy 4.0.1 (GPL-3.0), PyCBC 2.11.0 (GPL-3.0), bilby 2.8.2 (MIT) — https://gwosc.org/software/

OUTCOME: Declined budget_exhausted
DETAIL: solver budget spent
NO PROOF WAS JUDGED — the formalizer rejected the solver's answer:
The solution contains substantive BayesWave errors. First, Table 2 for Blip glitch 1 has 19/21—not 17/21—pre-subtraction 90%-interval successes: chirp mass 5/7, mass ratio 7/7, and effective spin 7/7; post-subtraction is 21/21. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) Second, the row incorrectly combines “BayesWave Joint / CBC+Glitch.” The reported injection study used the model-agnostic Joint configuration, while CBC+Glitch is a distinct configuration discussed as producing a less severe bias in the difficult overlap case; therefore the Joint results, including the failure below 0.75 overlap, cannot be attributed indiscriminately to CBC+Glitch. ([arxiv.org](https://arxiv.org/html/2311.09159v2)) The solution should separate those configurations and correct the coverage count.
TOOL ACTIVITY: formalizer:web_search x9, solver:web_search x21
LITERATURE SEARCH (engine-witnessed, openalex, as of 2026-09-02; total matches per query, queries broad -> narrow):
- CLAIM: What parameter-estimation bias does each glitch-subtraction method introduce in gravitational-wave parameter estimation, measured as the offset between injected and recovered parameters on public LIGO-Virgo-KAGRA data?  [UNCALIBRATED: no query matched anything, so this says nothing about absence]
    FAILED hits: "gravitational wave" AND glitch AND "parameter estimation"
    FAILED hits: "gravitational wave" AND ("glitch subtraction" OR "glitch mitigation") AND "parameter estimation"
    FAILED hits: "gravitational wave" AND ("glitch subtraction" OR "glitch mitigation") AND (injection OR injected) AND (recovery OR recovered)
    FAILED hits: "gravitational wave" AND glitch AND "parameter bias" AND injection
    FAILED hits: "gravitational wave" AND glitch AND "parameter estimation" AND (coverage OR "credible interval")
    FAILED hits: LIGO AND glitch AND "parameter estimation" AND "public data"
    FAILED hits: GWOSC AND glitch AND "parameter estimation" AND injection
    FAILED hits: BayesWave AND glitch AND "parameter estimation" AND injection
    FAILED hits: BayesWave AND Joint AND glitch AND "parameter estimation"
    FAILED hits: BayesWave AND "CBC+Glitch" AND "parameter estimation"
    FAILED hits: gwsubtract AND glitch AND "parameter estimation"
    FAILED hits: ("gating" OR inpainting) AND glitch AND "gravitational wave" AND "parameter estimation"
    FAILED hits: ("glitch subtraction" OR "glitch mitigation") AND "injected and recovered" AND (LIGO OR Virgo OR KAGRA)
    FAILED hits: ("glitch subtraction" OR "glitch mitigation") AND (bias OR offset) AND injection AND recovered AND (GWOSC OR "public data")
STATS: 3 calls, 166.7s

reasoning

{
 "summary": [
  "**Clarifying research needs**\n\nI’m realizing that I need an exact JSON schema. It seems there’s no proof leading to a verdict, and some literature appears uncalibrated. The evidence suggests existing studies, yet I must question the methods used and the public data availability. It’s uncertain whether a comprehensive analysis across methods has been published. I should review specific items, verify the literature, define methods clearly, check BayesWave configurations, and ensure public datasets are accessible. I’ll aim for a grounded summary."
 ]
}

output

{
 "verdict": "maybe",
 "explanation": "The run did not establish openness: no proof was judged, every literature-search rung failed, and the cited BayesWave evidence already contains at least a partial injection study whose configurations and coverage results were misreported.",
 "review_items": [
  "Run a calibrated literature search to determine whether a published study already compares injected-versus-recovered parameter offsets across glitch-subtraction methods on public LVK data.",
  "Check whether the required public GWOSC strain, glitch examples, and injection metadata exist for a reproducible comparison.",
  "Define which methods and configurations count as distinct, explicitly separating BayesWave Joint from CBC+Glitch.",
  "Correctly extract the BayesWave results, including the 19/21 pre-subtraction and 21/21 post-subtraction 90%-interval successes for Blip glitch 1, and determine whether they already answer part of the proposal.",
  "Specify a common bias statistic and parameter set that permit meaningful comparisons across methods and datasets."
 ]
}
setup osf · gpt-5.6-sol @ medium · {'requests': 2, 'input_tokens': 180257, 'cached_tokens': 0, 'output_tokens': 6801, 'reasoning_tokens': 4854} · json record