Methodology
How GrivavaLAB tests correlations between space weather and human-health indicators — transparently, reproducibly, and with honest treatment of null results.
Evidence-firstOpen dataReproducible
1. Research Question
Heliobiology (Chizhevsky tradition) proposes that solar and geomagnetic activity correlate with human physiology and behaviour. We test this hypothesis at population scale using: (a) public space-weather archives, and (b) anonymised aggregate search-interest data (e.g. Google Trends) as behavioural/health proxies.
Key stance: a statistically significant correlation is not proof of causation. We report effect sizes, control for multiple comparisons, and run a robustness battery before claiming any finding.
2. Data Sources
- NASA OMNIWeb — solar wind (speed, density, temperature), IMF (B, Bz), geomagnetic indices. Hourly, 1963–present.
- NOAA SWPC — Kp, Ap, Dst planetary indices. 1932–present, 3-hourly.
- NASA DONKI — CME catalogue, solar flares, proton events. 2010–present.
- GOES — X-ray flux, proton flux. 1998–present, 1-minute.
- SILSO — international sunspot number. 1749–present.
- DSCOVR / ACE (L1) — real-time solar wind at Lagrangian point 1.
- Google Trends — daily/weekly relative search interest for health-related terms (e.g. epilepsy, insomnia, migraine, cortisol).
3. Correlation Pipeline
- Data ingestion — pull both space-weather and search-interest series; keep units and time resolution.
- Temporal alignment — resample to a common cadence (daily), align on UTC timestamps.
- Detrending / deseasonalisation — STL decomposition (365-day seasonal period) to remove annual cycles and long-term trends that could confound correlations.
- Normalisation — z-score residuals before correlation.
- Lagged cross-correlation (CCF) — compute Pearson r for lags
−30…+30 days(space weather leading, coincident, or lagging the indicator). - Multiple-comparison control — Benjamini–Hochberg FDR correction at
α = 0.05across all lags × variables. - Surrogate testing — compare observed r against a null distribution from 1,000 phase-randomised / permutation surrogates.
4. Robustness Battery
Any candidate finding must survive all of these before we call it noteworthy:
- FDR-corrected significance — survives multiple-comparison control.
- Subperiod stability — the correlation holds across non-overlapping sub-samples (e.g. split halves).
- Block bootstrap — confidence interval on r excludes zero under serially-correlated resampling.
- Surrogate / placebo test — the signal is not reproduced with shuffled or phase-randomised surrogate series.
- Cross-variable check — the effect is specific to the hypothesised pair, not a general artefact of the data generation process.
Why this matters: our "laziness" study initially found r = +0.102 (FDR-significant) that collapsed under robustness testing — a textbook Simpson's-paradox case. This is exactly what the battery is designed to catch.
5. Reporting Rules
- We report effect sizes (correlation r, sample size n, p-value, FDR q-value), never just p < 0.05.
- We report null results as prominently as positive ones.
- We explicitly state that findings are associational, not causal.
- We never interpret space-weather values as predictors of individual health, mood, or behaviour.
- All pipelines and intermediate data are documented for independent verification.
6. Tools & Environment
Python 3, pandas, SciPy, statsmodels (STL), statsmodels.tsa.stattools (CCF), statsmodels.stats.multitest (BH-FDR). Reruns are deterministic given the same input data.
7. Limitations
- Google Trends data are relative, sampled, and normalised — effect sizes are approximate and confined to the platform's geography/tuning.
- Multiple comparisons across many terms × variables × lags inflate false-positive risk; controlled only by the FDR step.
- Confounders (seasonality, news cycles, economic stress) are only partially removed by STL detrending.
- Population aggregates cannot speak to individual physiology — any individual-level inference is out of scope by design.