OCO‑2 · OCO-3 · ERA5 · HGB v2
Click anywhere on India to select a point

Charts appear here after querying a point

◎ Choose a mode, then select a region
🗺️

Select a region to begin

Choose Point (click map + radius), City, or State — then set a date range and click Analyse.

🧠 Model Architecture & Inference Pipeline
HistGradientBoosting · 24 per-hour models · ERA5 BLH + OCO-2/3 satellite anchor · physics-informed smoothing
Data Flow Pipeline
🛰️
OCO-2/3
Satellite XCO₂
anchor @ hr 13
→
🌤️
ERA5 NetCDF
24h boundary
layer heights
→
⚙️
Feature Eng.
11 features
per hour
→
📐
StandardScaler
Per-hour
normalisation
→
🤖
HGB model_h
24 models
→ Δ CO₂
→
〰️
Smoothing
Gaussian blur
+ BLH physics
→
📈
Output
24h CO₂
profile (ppm)
Input Features (11)
#FeatureTypeDescription
1blhMeteoERA5 BLH at target hour (m)
2blh_ratioMeteoblh / daily mean BLH
3log_blhMeteolog(blh + 1) — log-scale mixing
4sin_hourTimesin(2π·h/24) — cyclic encoding
5cos_hourTimecos(2π·h/24) — cyclic encoding
6co2_13AnchorSatellite XCO₂ at hr 13 (ppm)
7blh_13AnchorERA5 BLH at hr 13 (m)
8sin_doyTimesin(2π·doy/365) — seasonal
9cos_doyTimecos(2π·doy/365) — seasonal
10seasonCalendar0=Winter 1=Spring 2=Summer 3=Autumn
11is_weekendCalendar0/1 — weekend emission flag
Model Details
🤖 Algorithm
HistGradientBoostingRegressor (sklearn) — histogram-based gradient boosting, fast on large datasets, native missing-value support.
Count24 models (one / hour)
🎯 Target Variable
Predicts Δ CO₂ = CO₂_h − CO₂_13. Final output = co2_13 + Δ. Anchoring to satellite prevents accumulative drift.
Unitsppm (deviation)
📐 Scaler
Per-hour StandardScaler fitted on named pd.DataFrame. Must receive column-named input — avoids sklearn UserWarning on wrong feature order.
TypeX-only (not target)
〰️ Smoothing
Gaussian blur σ=1.8h (circular wrap) + 35% BLH physics blend (CO₂ ∝ −BLH). Hard re-anchor at hr 13.
Alpha0.35 physics / 0.65 model
🔐 No Leakage
Temporal split by date (not row). Optuna tunes on validation only. Test set touched once. Zero future CO₂ leakage — only co2_13 anchor used.
Split70 / 15 / 15
⚡ Inference Speed
24 model.predict() calls with pre-loaded pkl files. BLH NetCDF read is the bottleneck (~0.3s). Total query response typically <1s.
BottleneckERA5 NetCDF I/O
Active hour models
Live Inference Simulation Watch the pipeline execute step-by-step for a sample query

// Click ▶ Run to simulate the full inference pipeline

0 / 24 hours