Methodology

How GrivavaLAB tests correlations between space weather and human-health indicators — transparently, reproducibly, and with honest treatment of null results.

Evidence-firstOpen dataReproducible

1. Research Question

Heliobiology (Chizhevsky tradition) proposes that solar and geomagnetic activity correlate with human physiology and behaviour. We test this hypothesis at population scale using: (a) public space-weather archives, and (b) anonymised aggregate search-interest data (e.g. Google Trends) as behavioural/health proxies.

Key stance: a statistically significant correlation is not proof of causation. We report effect sizes, control for multiple comparisons, and run a robustness battery before claiming any finding.

2. Data Sources

  • NASA OMNIWeb — solar wind (speed, density, temperature), IMF (B, Bz), geomagnetic indices. Hourly, 1963–present.
  • NOAA SWPC — Kp, Ap, Dst planetary indices. 1932–present, 3-hourly.
  • NASA DONKI — CME catalogue, solar flares, proton events. 2010–present.
  • GOES — X-ray flux, proton flux. 1998–present, 1-minute.
  • SILSO — international sunspot number. 1749–present.
  • DSCOVR / ACE (L1) — real-time solar wind at Lagrangian point 1.
  • Google Trends — daily/weekly relative search interest for health-related terms (e.g. epilepsy, insomnia, migraine, cortisol).

3. Correlation Pipeline

  1. Data ingestion — pull both space-weather and search-interest series; keep units and time resolution.
  2. Temporal alignment — resample to a common cadence (daily), align on UTC timestamps.
  3. Detrending / deseasonalisation — STL decomposition (365-day seasonal period) to remove annual cycles and long-term trends that could confound correlations.
  4. Normalisation — z-score residuals before correlation.
  5. Lagged cross-correlation (CCF) — compute Pearson r for lags −30…+30 days (space weather leading, coincident, or lagging the indicator).
  6. Multiple-comparison control — Benjamini–Hochberg FDR correction at α = 0.05 across all lags × variables.
  7. Surrogate testing — compare observed r against a null distribution from 1,000 phase-randomised / permutation surrogates.

4. Robustness Battery

Any candidate finding must survive all of these before we call it noteworthy:

  1. FDR-corrected significance — survives multiple-comparison control.
  2. Subperiod stability — the correlation holds across non-overlapping sub-samples (e.g. split halves).
  3. Block bootstrap — confidence interval on r excludes zero under serially-correlated resampling.
  4. Surrogate / placebo test — the signal is not reproduced with shuffled or phase-randomised surrogate series.
  5. Cross-variable check — the effect is specific to the hypothesised pair, not a general artefact of the data generation process.

Why this matters: our "laziness" study initially found r = +0.102 (FDR-significant) that collapsed under robustness testing — a textbook Simpson's-paradox case. This is exactly what the battery is designed to catch.

5. Reporting Rules

  • We report effect sizes (correlation r, sample size n, p-value, FDR q-value), never just p < 0.05.
  • We report null results as prominently as positive ones.
  • We explicitly state that findings are associational, not causal.
  • We never interpret space-weather values as predictors of individual health, mood, or behaviour.
  • All pipelines and intermediate data are documented for independent verification.

6. Tools & Environment

Python 3, pandas, SciPy, statsmodels (STL), statsmodels.tsa.stattools (CCF), statsmodels.stats.multitest (BH-FDR). Reruns are deterministic given the same input data.

7. Limitations

  • Google Trends data are relative, sampled, and normalised — effect sizes are approximate and confined to the platform's geography/tuning.
  • Multiple comparisons across many terms × variables × lags inflate false-positive risk; controlled only by the FDR step.
  • Confounders (seasonality, news cycles, economic stress) are only partially removed by STL detrending.
  • Population aggregates cannot speak to individual physiology — any individual-level inference is out of scope by design.

← Back to Research Log