Skip to content

Hermes Forecast Methodology: Falsifiable Predictions & Epistemic Resolution Oracles


1. The Epistemic Crisis of Threat Intelligence

Section titled “1. The Epistemic Crisis of Threat Intelligence”

Traditional cybersecurity reporting is dominated by rhetorical ambiguity:

  • “Threat actors may increasingly target edge gateway appliances in the coming months.”
  • “Organizations should expect continued exploitation of memory corruption vulnerabilities.”

These statements appear insightful, but they violate Karl Popper’s criterion of falsifiability:

  1. No Target Date: If no attack occurs in 6 months, the author argues the threat is “still evolving.” If an attack occurs after 3 years, they claim credit.
  2. No Quantified Probability: If an event has a 5% chance of occurring and happens, the author claims foresight. If an event has a 95% chance and fails to happen, they claim “the landscape shifted.”
  3. No External Oracle: The analyst acts as their own judge, deciding whether an event satisfies their own vague criteria.

The Hermes Forecast Methodology replaces rhetorical evasion with formal probabilistic forecasting, strict time horizons, and deterministic third-party adjudication.


Explore live predictions, inspect verified resolution proofs, and test counterfactual outcome simulations:

Mean Brier Score (BS)
0.1043
58% vs random baseline (0.25)
Resolution Track Record
5 / 13
4 TRUE 1 FALSE
Active In-Flight Forecasts
8
Horizons 14 to 60 days
Calibration Tier
HIGH
BS ≤ 0.15 (High calibration)
📊

Reliability Diagram & Empirical Calibration (Deciles)

PRINCIPLE P10: NO FAKE PRECISION

A well-calibrated forecast matches predicted probability with empirical realization frequency. The deliberate inclusion of failed forecasts (e.g. CVE-2025-3248) prevents hindsight cherry-picking and confirms epistemic integrity (Principle P5).

Probability Range Forecast Count Mean Forecast Empirical Outcome Calibration Delta Visual Balance
0.00 - 0.20 0 (0 resolved) — Pending —
0.20 - 0.40 0 (0 resolved) — Pending —
0.40 - 0.60 0 (0 resolved) — Pending —
0.60 - 0.80 7 (2 resolved) 73% 50% +0.23
0.80 - 1.00 6 (3 resolved) 87% 100% -0.13
🔮

Counterfactual 'What-If' Resolution Simulator

Simulate how active predictions would alter the global Brier Score in real-time
Simulated BS: 0.1043

3. Mathematical Foundations: The Brier Score & Calibration

Section titled “3. Mathematical Foundations: The Brier Score & Calibration”

To evaluate whether Hermes’s forecasts are genuinely calibrated or merely lucky guesses, we implement the Brier Score ($BS$), developed by Glenn W. Brier (1950):

For an individual prediction i, the Brier score measures the squared difference between the forecasted probability f_i and the binary empirical outcome o_i:

BS_i = (f_i - o_i)^2

Where:

  • f_i is the forecasted probability assigned at inception ($0.0 \le f_i \le 1.0$)
  • o_i is the binary empirical resolution ($o_i = 1$ if verified by oracle before cutoff; $o_i = 0$ if cutoff elapsed without event)

Across a corpus of N resolved predictions, the aggregate performance is:

BS = (1 / N) * SUM_{i=1}^N (f_i - o_i)^2

The Brier score is a strictly proper scoring rule: the forecaster minimizes their expected penalty if and only if they report their true honest belief. Overconfident forecasters who assign 99% to uncertain events suffer massive quadratic penalties ((0.99 - 0)^2 = 0.9801), while cautious under-forecasters suffer penalties for under-weighting high-probability events.

To benchmark Hermes against a baseline random forecaster (assigning $P = 0.50$ to every event, resulting in a baseline BS_ref = 0.25):

BSS = 1 - (BS / BS_ref) = 1 - (BS / 0.25)
  • A positive $BSS$ ($> 0%$) indicates genuine predictive edge over uninformed guessing.
  • A score of $+58%$ (Hermes operational baseline at $BS \approx 0.1043$) demonstrates high empirical calibration.

4. Epistemic Oracles & Ground-Truth Verification

Section titled “4. Epistemic Oracles & Ground-Truth Verification”

A prediction cannot be adjudicated by the analyst who created it. Hermes delegates resolution to Deterministic Oracles:

+-------------------------------------------------------------------------------+
| ORACLE TAXONOMY |
+-------------------------------------------------------------------------------+
| Oracle Type | Authoritative Source | Adjudication Rule |
+-------------------------------+-------------------------+---------------------+
| cisa_kev_inclusion | CISA KEV JSON Feed | Exact CVE ID match |
| verified_public_weaponization | GitHub / Exploit-DB | Working PoC exploit |
| greynoise_mass_exploitation | GreyNoise Sensor Array | Threshold IP scans |
| shadowserver_telemetry | Shadowserver Foundation | Attack packet logs |
| cert_national_advisory | ANSSI / CISA / BSI | Formal public alert |
+-------------------------------------------------------------------------------+
  1. Time Horizon Expiration: If the target date (e.g., 2026-10-15T00:00:00Z) elapses without oracle verification, the prediction status is automatically set to RESOLVED_FALSE with $o_i = 0$.
  2. Early Verification: If the oracle verifies the event on Day 3 of a 30-day window, the prediction resolves immediately as RESOLVED_TRUE with $o_i = 1$. The horizon date remains recorded for historical auditability.
  3. No Post-Hoc Adjustments (Principle P5): Neither probability, horizon, nor criterion may be adjusted after publication.

5. Anti-Hindsight Epistemic Honesty: Why We Publish False Predictions

Section titled “5. Anti-Hindsight Epistemic Honesty: Why We Publish False Predictions”

In human cognition, hindsight bias makes past events seem predictable (“we always knew that would happen”). Unprincipled intelligence platforms hide their incorrect forecasts to present an illusion of 100% foresight.

Hermes deliberately displays predictions that failed to resolve true. For example:

  • PRED-003 (CVE-2025-3248 Langflow Memory Corruption): Forecasted at $P = 0.65$ for mass opportunistic internet scanning within 30 days. When the window expired, GreyNoise logged only 11 benign research scans; exploit fragility prevented botnet adoption.
  • The prediction was permanently logged as RESOLVED_FALSE with $BS = 0.4225$.

This failure serves two critical architectural functions:

  1. Model Recalibration: It revealed that high CVSS severity in Python C-extensions does not correlate with fast weaponization if memory alignment prerequisites cause target server panics.
  2. Mathematical Authenticity: A model that reports a mean Brier Score of 0.1043 while openly auditing its 0.4225 errors possesses far greater institutional credibility than an opaque feed claiming zero errors.

6. Strategic Integration in the Hermes Moat

Section titled “6. Strategic Integration in the Hermes Moat”

Hermes Forecast represents Phase 2.0 of the Hermes Intelligence Roadmap:

V1.0 Foundation (Ontology & Evidence Graph)
↓
V1.1 Hermes Observatory (Signal Ingestion & Threat Funnel)
↓
V1.2 Risk Trajectory (Dynamic Velocity & Acceleration Modeling)
↓
V1.3 Historical Intelligence (Daily Snapshots & Time Machine)
↓
V1.4 Vulnerability Genome (Modular Loci & Genetic Distance)
↓
V1.5 Hermes Autopsy (Reverse Forensic Post-Mortems)
↓
V2.0 Hermes Forecast (Falsifiable Predictions & Resolution Oracles)

By connecting the Vulnerability Genome (identifying architectural bug patterns) to Hermes Forecast (predicting weaponization likelihoods) and Hermes Autopsy (forensically verifying the failure points when breaches occur), Hermes establishes an integrated, self-calibrating cyber intelligence platform.