Hermes Forecast Methodology: Falsifiable Predictions & Epistemic Resolution Oracles
1. The Epistemic Crisis of Threat Intelligence
Section titled “1. The Epistemic Crisis of Threat Intelligence”Traditional cybersecurity reporting is dominated by rhetorical ambiguity:
- “Threat actors may increasingly target edge gateway appliances in the coming months.”
- “Organizations should expect continued exploitation of memory corruption vulnerabilities.”
These statements appear insightful, but they violate Karl Popper’s criterion of falsifiability:
- No Target Date: If no attack occurs in 6 months, the author argues the threat is “still evolving.” If an attack occurs after 3 years, they claim credit.
- No Quantified Probability: If an event has a 5% chance of occurring and happens, the author claims foresight. If an event has a 95% chance and fails to happen, they claim “the landscape shifted.”
- No External Oracle: The analyst acts as their own judge, deciding whether an event satisfies their own vague criteria.
The Hermes Forecast Methodology replaces rhetorical evasion with formal probabilistic forecasting, strict time horizons, and deterministic third-party adjudication.
2. Interactive Forecast Workbench
Section titled “2. Interactive Forecast Workbench”Explore live predictions, inspect verified resolution proofs, and test counterfactual outcome simulations:
Reliability Diagram & Empirical Calibration (Deciles)
A well-calibrated forecast matches predicted probability with empirical realization frequency. The deliberate inclusion of failed forecasts (e.g. CVE-2025-3248) prevents hindsight cherry-picking and confirms epistemic integrity (Principle P5).
| Probability Range | Forecast Count | Mean Forecast | Empirical Outcome | Calibration Delta | Visual Balance |
|---|---|---|---|---|---|
| 0.00 - 0.20 | 0 (0 resolved) | — | Pending | — | |
| 0.20 - 0.40 | 0 (0 resolved) | — | Pending | — | |
| 0.40 - 0.60 | 0 (0 resolved) | — | Pending | — | |
| 0.60 - 0.80 | 7 (2 resolved) | 73% | 50% | +0.23 | |
| 0.80 - 1.00 | 6 (3 resolved) | 87% | 100% | -0.13 |
Counterfactual 'What-If' Resolution Simulator
Simulate how active predictions would alter the global Brier Score in real-time3. Mathematical Foundations: The Brier Score & Calibration
Section titled “3. Mathematical Foundations: The Brier Score & Calibration”To evaluate whether Hermes’s forecasts are genuinely calibrated or merely lucky guesses, we implement the Brier Score ($BS$), developed by Glenn W. Brier (1950):
Individual Prediction Score
Section titled “Individual Prediction Score”For an individual prediction i, the Brier score measures the squared difference between the forecasted probability f_i and the binary empirical outcome o_i:
BS_i = (f_i - o_i)^2Where:
f_iis the forecasted probability assigned at inception ($0.0 \le f_i \le 1.0$)o_iis the binary empirical resolution ($o_i = 1$ if verified by oracle before cutoff; $o_i = 0$ if cutoff elapsed without event)
Global Mean Brier Score
Section titled “Global Mean Brier Score”Across a corpus of N resolved predictions, the aggregate performance is:
BS = (1 / N) * SUM_{i=1}^N (f_i - o_i)^2The Brier score is a strictly proper scoring rule: the forecaster minimizes their expected penalty if and only if they report their true honest belief. Overconfident forecasters who assign 99% to uncertain events suffer massive quadratic penalties ((0.99 - 0)^2 = 0.9801), while cautious under-forecasters suffer penalties for under-weighting high-probability events.
Brier Skill Score ($BSS$)
Section titled “Brier Skill Score ($BSS$)”To benchmark Hermes against a baseline random forecaster (assigning $P = 0.50$ to every event, resulting in a baseline BS_ref = 0.25):
BSS = 1 - (BS / BS_ref) = 1 - (BS / 0.25)- A positive $BSS$ ($> 0%$) indicates genuine predictive edge over uninformed guessing.
- A score of $+58%$ (Hermes operational baseline at $BS \approx 0.1043$) demonstrates high empirical calibration.
4. Epistemic Oracles & Ground-Truth Verification
Section titled “4. Epistemic Oracles & Ground-Truth Verification”A prediction cannot be adjudicated by the analyst who created it. Hermes delegates resolution to Deterministic Oracles:
+-------------------------------------------------------------------------------+| ORACLE TAXONOMY |+-------------------------------------------------------------------------------+| Oracle Type | Authoritative Source | Adjudication Rule |+-------------------------------+-------------------------+---------------------+| cisa_kev_inclusion | CISA KEV JSON Feed | Exact CVE ID match || verified_public_weaponization | GitHub / Exploit-DB | Working PoC exploit || greynoise_mass_exploitation | GreyNoise Sensor Array | Threshold IP scans || shadowserver_telemetry | Shadowserver Foundation | Attack packet logs || cert_national_advisory | ANSSI / CISA / BSI | Formal public alert |+-------------------------------------------------------------------------------+Adjudication Rules
Section titled “Adjudication Rules”- Time Horizon Expiration: If the target date (e.g.,
2026-10-15T00:00:00Z) elapses without oracle verification, the prediction status is automatically set toRESOLVED_FALSEwith $o_i = 0$. - Early Verification: If the oracle verifies the event on Day 3 of a 30-day window, the prediction resolves immediately as
RESOLVED_TRUEwith $o_i = 1$. The horizon date remains recorded for historical auditability. - No Post-Hoc Adjustments (Principle P5): Neither probability, horizon, nor criterion may be adjusted after publication.
5. Anti-Hindsight Epistemic Honesty: Why We Publish False Predictions
Section titled “5. Anti-Hindsight Epistemic Honesty: Why We Publish False Predictions”In human cognition, hindsight bias makes past events seem predictable (“we always knew that would happen”). Unprincipled intelligence platforms hide their incorrect forecasts to present an illusion of 100% foresight.
Hermes deliberately displays predictions that failed to resolve true. For example:
PRED-003(CVE-2025-3248 Langflow Memory Corruption): Forecasted at $P = 0.65$ for mass opportunistic internet scanning within 30 days. When the window expired, GreyNoise logged only 11 benign research scans; exploit fragility prevented botnet adoption.- The prediction was permanently logged as
RESOLVED_FALSEwith $BS = 0.4225$.
This failure serves two critical architectural functions:
- Model Recalibration: It revealed that high CVSS severity in Python C-extensions does not correlate with fast weaponization if memory alignment prerequisites cause target server panics.
- Mathematical Authenticity: A model that reports a mean Brier Score of
0.1043while openly auditing its 0.4225 errors possesses far greater institutional credibility than an opaque feed claiming zero errors.
6. Strategic Integration in the Hermes Moat
Section titled “6. Strategic Integration in the Hermes Moat”Hermes Forecast represents Phase 2.0 of the Hermes Intelligence Roadmap:
V1.0 Foundation (Ontology & Evidence Graph) ↓V1.1 Hermes Observatory (Signal Ingestion & Threat Funnel) ↓V1.2 Risk Trajectory (Dynamic Velocity & Acceleration Modeling) ↓V1.3 Historical Intelligence (Daily Snapshots & Time Machine) ↓V1.4 Vulnerability Genome (Modular Loci & Genetic Distance) ↓V1.5 Hermes Autopsy (Reverse Forensic Post-Mortems) ↓V2.0 Hermes Forecast (Falsifiable Predictions & Resolution Oracles)By connecting the Vulnerability Genome (identifying architectural bug patterns) to Hermes Forecast (predicting weaponization likelihoods) and Hermes Autopsy (forensically verifying the failure points when breaches occur), Hermes establishes an integrated, self-calibrating cyber intelligence platform.