← COVER

Can We Ever Understand a Planet?

COSMICS · PREFACE

Can We Ever Build Systems
That Understand a Planet?

On the distance between measuring Earth and understanding it.

EARTHVISION LAB · ~28 MIN READ

The heliacal rising of Sirius, its first visible appearance above the eastern horizon after weeks of solar conjunction, fell within days of the Nile's annual inundation. Egyptian priests used that single stellar event to anchor a 365-day calendar and time agricultural cycles for a river they never gauged.

A comparable feat, arrived at independently, stands in Jaipur. The Samrat Yantra, a stone gnomon over twenty metres tall, reads solar time to roughly two seconds and tracked declination and hour angle well enough to fix eclipses and monsoon-relevant solstices. No lens, no electronics, no orbit. Just geometry, repetition, and a few centuries of correction.

Both systems solved a forecasting problem with zero instrumentation in the modern sense: a star position and a shadow's length, converted into a decision about when to plant. Multispectral imagers, synthetic aperture radar, and gravimetric satellites now return more measurements of this planet in a single orbit than either civilisation collected in a generation. Whether that volume of measurement has converted into comparable forecasting skill is a separate, open question, and it is the one this note works through: not as a thought experiment, but as a distinction between two capabilities that get routinely conflated: observation, and understanding.

01

Observation: a decade of hardware growing up

A decade ago, a functioning Earth-observation payload meant a multi-year procurement: a custom bus, a dedicated launch slot, a licensed ground station, and a failure budget large enough that a single bad separation event could end the mission and the company behind it. Cost to low Earth orbit sat well above ten thousand dollars per kilogram for most operators.

Reusable first stages compressed that figure by roughly an order of magnitude and shortened launch cadence from years to weeks. In parallel, the CubeSat standard (3U, 6U, and now ESPA-class rideshare payloads) let sensor packages shrink from bus-sized instruments to units small enough to fly as secondary payloads on someone else's mission, splitting fixed launch cost across dozens of operators on a single manifest.

Sensor design moved with it. Multispectral imagers dropped from a handful of broad bands to a dozen or more, narrower ones; synthetic aperture radar, once the preserve of national programmes, now flies on constellations built by companies with headcounts in the low hundreds. Revisit cadence for a given point on Earth went from weeks to, for some constellations, several times a day. Ground segments moved off dedicated antenna farms and onto cloud downlink networks, so a new operator can license capacity instead of building infrastructure.

The compounded effect is observation priced and delivered as a utility: a burn scar, a flood extent, a construction footprint, resolvable within hours at a cost per square kilometre that keeps falling. That is real, hard-won infrastructure. It resolves where and what happened with a precision no prior instrument set could match. It was not built to resolve what happens next, and higher resolution does not close that gap on its own.

Launch cadence and per-kilogram cost, falling together
02

Understanding: what it would actually take

Detection and forecasting are different statistical problems. Detection asks whether a pattern already present in the data matches a known signature. Forecasting asks whether a combination of variables, none individually anomalous, is converging toward a state that has not occurred yet. Most operational Earth-observation products, active-fire detections included, are built to solve the first problem.

A three-hour lead time on a wildfire is a detection problem with a tight service level: fuse thermal-anomaly returns from a VIIRS- or MODIS-class sensor with near-real-time wind direction and terrain slope, fast enough to outrun the fire's own spread rate. A three-day lead time is a different class of problem. It requires reasoning jointly over fuel moisture content, vapour-pressure deficit, a fire-weather index such as Fosberg or the Canadian FWI, antecedent drought conditions, and human-activity proxies like nightlight anomalies or road proximity, in a region where no ignition has happened and no sensor has flagged anything at all.

That is the actual requirement behind the word: a system that holds probabilistic estimates across variables nobody engineered to be read together, and outputs a forecast with calibrated uncertainty rather than a confirmed report. Most current architectures are trained and scored the opposite way, on single-variable accuracy against ground truth that, by definition, already happened.

A model scored only on detections learns to recognise fire. It is never asked to anticipate one.

Three structural reasons explain the gap. Instrument heterogeneity: missions built by different agencies and companies rarely share a spatial grid, revisit schedule, or spectral calibration, so fusion requires resampling that discards information before a model ever sees it. Evaluation incentive: a false detection is cheap to audit and explain, while an unrealised forecast carries reputational cost, so benchmarks reward recall on known events over calibrated probability on unknown ones. Scope: the discipline optimised for coverage and resolution because those were the tractable engineering targets, and largely stopped there.

Closing it requires architectures with temporal state, models that carry a running estimate of a region's condition instead of scoring each pass independently, evaluated on the sharpness and calibration of early probabilistic warnings rather than on final-detection recall alone, and trained across domains jointly, so a drought index, a grid-load anomaly, and a vegetation-moisture trend inform the same estimate instead of feeding three unconnected pipelines.

Detection at ignition, versus forecast at the fuel stage
03

Reading a field without a satellite

NDVI, the normalised difference between near-infrared and red reflectance, correlates with canopy chlorophyll density and has been the default proxy for crop vigour since Landsat data made it computable at scale in the 1970s. It saturates in dense canopy, though, which is why EVI and chlorophyll-sensitive red-edge indices were added later to hold sensitivity through peak growth stages.

None of these indices measure yield directly, only a physiological correlate of it, several weeks removed from harvest. That correlate was, in effect, read manually long before it had a spectral name: leaf angle, soil crust texture, and insect behaviour, tracked by growers whose only instrument was memory of the previous decade's failures.

L-band radar, which penetrates canopy to read soil dielectric properties, and SMAP-derived root-zone moisture estimates now add a variable that colour-based indices cannot see at all: water available below the surface, days before canopy stress becomes visible. The open problem is not replacing generational field knowledge with these measurements. It is building a model that treats a smallholder's seasonal judgment and a soil-moisture retrieval as two noisy estimates of the same latent variable, and fuses them accordingly.

04

The grid, read from orbit

Every sensor, ground station, and inference pipeline described so far depends on a grid staying within its operating tolerances. That tolerance is a live balancing act across generation, transmission loss, and demand, and it grows harder to hold as intermittent sources take a larger share of supply. Tracking that balance from the ground means waiting on utility filings and quarterly disclosures; tracking it from orbit means reading the physical activity that precedes those filings.

That physical activity has spectral signatures of its own. Shortwave-infrared sensors tuned for gas flaring detect combustion distinct from wildfire, and methane-sensitive instruments such as TROPOMI resolve plume concentrations down to tens of parts per billion, catching leaks before they appear in any reported inventory. VIIRS nightfire products, originally built to flag combustion, now double as a proxy for tracking flaring intensity by field and by operator, month over month.

The same sensors that catch a flare at the scale of a single well can be aggregated to a coarser signal: nightlight radiance, summed over a grid cell and tracked over time, correlates with regional electricity consumption closely enough that several central banks use it as a nowcast for economic activity where official statistics lag by months. Read together, flaring, nightlight, and plant-level classification turn the energy transition from a policy claim into something closer to a measured quantity, resolvable from orbit faster than most reporting cycles disclose it.

05

Feedback, not a single signal

Sea ice has an albedo near 0.6 to 0.9; open ocean water sits closer to 0.06. Each square kilometre of ice lost exposes water that absorbs far more incoming shortwave radiation, warming the adjacent water column and accelerating further melt. This is the ice-albedo feedback, one of a short list of processes climate scientists classify as self-reinforcing rather than linear.

GRACE and GRACE-FO satellites measure ice mass loss not by imaging but by tracking minute changes in gravitational field between twin spacecraft, a method sensitive enough to resolve Antarctic and Greenland mass balance to within tens of gigatonnes per year. That number, on its own, is a single reading. Its consequence three domains away, in monsoon timing over the Indian subcontinent or in fishery productivity along a cold-water coastline, requires a coupled ocean-atmosphere-cryosphere model to trace, not a satellite pass.

Permafrost carries a second, slower-acting feedback: as it thaws, microbial decomposition releases stored carbon and methane, which adds forcing that accelerates further thaw. Solstice and flood-cycle tracking by early astronomers were, functionally, single-variable models of one recurring feedback loop. The unresolved engineering task is a model wide enough in scope to carry several such loops simultaneously without losing the coupling between them.

06

Models that generalise past their training set

Foundation models for Earth observation, pretrained on large unlabelled multispectral and SAR archives through self-supervised objectives, now transfer to downstream tasks such as land-cover classification or flood mapping with a fraction of the labelled data earlier architectures required. That is a genuine advance in sample efficiency. It is not yet evidence of cross-domain reasoning.

The harder benchmark is out-of-distribution generalisation: a model trained on documented fire regimes correctly assessing risk in a region and fuel type it has never seen, the way a wayfinding navigator in the Pacific could read an unfamiliar ocean swell using dead-reckoning principles built from thousands of previously observed ones, without a chart.

That comparison sets a concrete bar rather than a rhetorical one. The navigator's method generalised because it encoded causal structure, how swell direction relates to distant weather systems, not just correlation with previously seen wave patterns. Architectures that fuse soil moisture, sea-surface temperature, grid load, and deforestation signal into a single spatiotemporal representation are a step toward that kind of structural encoding. Whether they generalise the way the navigator did remains an empirical question, not yet a settled one.

Cross-domain signals, fused before inference
07

Planetary intelligence, defined narrowly

Space systems, agriculture, energy, and climate science each developed as separate fields, with separate conferences, funding lines, and modelling conventions. Physically, they are coupled subsystems of one planet: a grid failure changes irrigation timing, a drought changes fire risk, a fire changes the carbon and albedo terms that feed back into the climate model measuring the drought.

Planetary intelligence, used here in a specific sense, names a closed-loop architecture: sensing, probabilistic modelling, and decision support coupled tightly enough that a change detected in one subsystem updates the operating estimate in the others, on a timescale short enough to inform a decision before the underlying event resolves itself.

No system meeting that definition currently exists at planetary scale. Component pieces do: gravimetric ice-mass tracking, methane-plume detection, soil-moisture retrieval, fire-weather indices, self-supervised Earth-observation models. The engineering work ahead is coupling, not invention: building the shared data model, the calibrated uncertainty estimates, and the cross-domain training objective that let these components inform one estimate instead of returning eight separate scores.