GEOINT · Spatial analytics · Methodology architecture
CBO–TGRA
Parcel Intelligence Under Uncertainty
Geospatial intelligence methodology integrating heterogeneous ownership, parcel, infrastructure, and geometric evidence into a calibrated, traceable prioritization of parcels near U.S. critical infrastructure.
01Operational question
The National Geospatial-Intelligence Agency required a method to identify parcels with foreign ownership or foreign beneficial ownership near U.S. critical infrastructure and to prioritize them for investigation. The decision behind the requirement: where should limited counterintelligence and facility-security analytical capacity be directed first, and with what stated confidence?
The government's starting point was a static CFIUS/MIRTA buffer. It produced a list without confidence, without geometry, and without a trail from priority back to evidence. I decomposed the requirement from first principles into five analytical questions, and the architecture followed that decomposition.
- Which parcels are relevant to the protected infrastructure, and by what rule?
- Who plausibly owns or controls each parcel?
- How reliable is the ownership evidence, and where does it conflict?
- What physical opportunity does the parcel's geometry create relative to the facility?
- How should those findings determine investigative priority for a given mission?
02Why the problem was difficult
Each question failed under the conventional approach for a different reason.
- Fragmented, heterogeneous data. More than 42 federal, regulatory, commercial, open-source, and geospatial sources, each with its own schema, identifiers, update cycle, geographic precision, evidentiary authority, completeness, and temporal behavior.
- Uncertain and adversarial evidence. Beneficial-ownership information arrives incomplete, contradictory, and at times deliberately obscured through layered corporate structures. Forcing it into an owned/not-owned determination manufactures certainty the evidence cannot support.
- Entity ambiguity. Whether records from different systems refer to the same person, company, address, or parcel is itself an inference with error, and it sits underneath every ownership claim.
- Spatial disagreement and geometric uncertainty. Terrain, observer and target positions, and sensor heights all carry error. A deterministic viewshed reports visible or not visible with more precision than the inputs justify.
- Conflated questions. Proximity screening blends two different claims, who controls the parcel and what the parcel physically enables, so a single score could not be interpreted, challenged, or recalibrated.
- Reproducibility. A federal analytic environment requires that any priority be reconstructed later with the information available at the time.
03My role
I originated Confidence-Based Ownership Assessment with Targeted Geometric Risk Analysis and served as methodology lead and principal architect under the USC–NGA Cooperative Research and Development Agreement.
- I architected the four-stage analytical methodology, the evidence model, the uncertainty framework, the validation approach, and the reproducibility controls.
- I designed the scoring and data architecture, the pipeline interfaces, the governance model, the analyst workflow, and the analyst-facing interface concept.
- I built representative KNIME, Python, ArcPy, and Esri geoprocessing components and Figma interface concepts, and authored the full technical specification.
- I briefed analysts, scientists, and NGA's Director of Counterintelligence at the abstraction level each audience needed, and I scoped the Washington, DC validation pilot on Regrid Premium parcel data that is now under way.
04Data architecture
Integration chain
Every source moves through the same chain, which is what makes 42 sources one system.
sourceingestionnormalizationcanonical representationentity or spatial relationshipevidence familyanalytical stageoutput
Positional and data-quality controls are explicit: CE95 for positional accuracy, ISO 19157 quality elements for completeness, logical consistency, and thematic accuracy, and QA gates on positional accuracy, schema conformity, attribute completeness, duplication, provenance completeness, and stage completeness.
05Methodology
Stage 0 · Expanded spatial screening
The first stage defines the analytical universe and governs the cost and coverage of every downstream stage. The static statutory buffer 𝔅 is retained and extended into 𝒰 = 𝔅 ∪ 𝔓 ∪ 𝔖: a shadow buffer 𝔖 that absorbs positional uncertainty, and four physics-informed exploitation envelopes 𝔓 computed against terrain and the road network for visual, radio-frequency, mobility, and unmanned-aircraft pathways. The envelope takes the farthest modality at every azimuth. Tipping and cueing is therefore built into the architecture: inexpensive broad screening first, expensive analysis only where evidence justifies it. The interactive explorer below computes the entire construction live.
Stage 1 · Ownership Confidence Score
Ownership attribution is an evidence problem, not a geometric one, so Stage 1 produces a calibrated Ownership Confidence Score instead of a binary determination.
source evidenceentity resolutionreliability assessmentprobabilistic and evidential combinationownership confidence
- Bayesian belief networks
- Conditional relationships among evidence and hypotheses let prior beliefs update as evidence arrives: prior evidence → conditional evidence → posterior ownership confidence. New evidence changes the posterior without rebuilding the methodology.
- Noisy-OR indicator families
- Several partially independent indicators can each support the same hypothesis without every combination being modeled by hand, which keeps heterogeneous ownership indicators tractable.
- Dempster–Shafer with Yager combination
- Where sources conflict, the calculus represents support, disbelief, and uncommitted uncertainty explicitly. Yager's rule assigns unresolved conflict to uncertainty instead of forcing it toward one conclusion, so disagreement between sources stays visible as an analytic signal.
- Three-tier entity resolution
- Obvious matches resolve deterministically on strong identifiers; ambiguous matches resolve through Fellegi–Sunter probabilistic linkage over names, addresses, jurisdiction, identifiers, and relationships; unresolved matches escalate to analyst adjudication, and the adjudicated outcome is committed back so the resolution is reused instead of rediscovered.
Stage 2 · Geometric Risk Score
Stage 2 asks a different question: what physical opportunity could this parcel create relative to the protected infrastructure? Four modalities run independently enough that an analyst can see which physical characteristic produced the assessment.
- Monte Carlo viewshed
- Define the uncertain parameters, sample plausible values, rerun the line-of-sight calculation repeatedly, and report the proportion of draws in which the sight line clears. The result is not visible or not visible; it is visibility supported across a stated share of plausible geometric conditions.
- Network chokepoints
- The road and access network becomes a graph; edge-betweenness centrality identifies segments that lie on a disproportionate share of shortest paths. Two parcels equally close to an installation can differ sharply in network significance.
- UAS feasibility geometry
- Parcel position, distance, terrain, elevation, and line of sight combine into a narrow physical question: does the spatial relationship create a plausible opportunity that deserves analyst attention? Geography is never used to infer intent.
- Change detection
- Observations through time separate significant physical change, a new structure, cleared ground, an altered access pattern, from normal parcel history and sensor variability. Change cues review; it is never treated as suspicious by itself.
Stage 3 · Terminal Fusion
Ownership Confidence Score, Geometric Risk Score, and mission weighting produce the Parcel Priority Score. Inference and decision remain separate: mission weighting changes what an analyst looks at first without retroactively changing the underlying evidence. A counterintelligence mission and a facility-security mission weight the same evidence differently and the evidence is unchanged.
Ownership confidence answers what the evidence supports about control. Geometric risk answers what the physical environment enables. Priority answers where limited analytical resources go first.
06Validation
Validation was designed into the methodology from the beginning under one principle: every consequential output needs a corresponding way to determine whether it deserves trust. Ownership inference, terrain geometry, and human classification make different kinds of claims and require different evidence.
| Method | What it establishes |
|---|---|
| Probability calibration | Whether stated confidence tiers correspond to empirical ownership outcomes; a 70 percent tier that is right 40 percent of the time is miscalibrated |
| False-positive and false-negative analysis | Where the methodology accepts or misses cases in ways that affect prioritization |
| Sensitivity analysis | Which assumptions and parameters dominate the result, and where rankings are fragile under small changes |
| Inter-rater reliability | Whether analysts applying the same rubric reach reproducible judgments |
| Spatial validation | Whether geometric outputs remain consistent with reference observations and CE95 positional accuracy |
| Model-appropriate attribution | Which features drove a result and how it connects to evidence; SHAP where technically valid |
| Conflict and stability monitoring | Whether source disagreement, model updates, or threshold changes materially alter priority |
Validation pilot
The Washington, DC pilot on Regrid Premium parcel data characterizes how the methodology behaves against a real parcel inventory before any threshold is treated as production performance. Thresholds and expected performance are design targets to be tested, not completed results.
- Score distributions and parcel volumes across review tiers
- Analyst workload created by proposed thresholds
- False-positive patterns and threshold-calibration needs
- Result stability under parameter change
- Discrimination contributed by each geometric modality
- Data-quality and source-coverage gaps
specificationprototypepilotempirical behaviorcalibrationrefinement
07Output
The four-stage architecture is drawn at full width in section 11: 42+ sources feed Stage 0 screening; the candidate universe flows in parallel to the Ownership Confidence Score and the Geometric Risk Score; Terminal Fusion combines them with mission weighting into the Parcel Priority Score, over a governance rail that runs beneath every stage.
Analyst-facing product
The output is not a ranked list. The interface concept answers, for any parcel, why it was prioritized, what ownership evidence contributed, which geometric modalities contributed, what confidence or uncertainty remains, what evidence needs review, and what supporting source data exists. Posterior, uncertainty, and conflict are retained as fields rather than collapsed into a score.
priorityfusion logicstage scoremodality or modelparametersevidencesource
Every priority reconstructs backward through that chain, so a challenge lands at the correct layer instead of on the whole system.
Delivered artifacts
- Full technical specification: mission decomposition, architecture, formal methods, data sources, schemas, interfaces, formulas, entity resolution, geospatial workflows, governance, QA, audit requirements, uncertainty, validation, pilot structure, analyst procedures, stage handoffs
- Representative KNIME ingestion workflows and Python, ArcPy, and Esri geoprocessing components: parcel screening, proximity, spatial joins, entity-resolution scaffolding, network accessibility, terrain analysis, visibility, change detection, scoring, QA
- Figma concepts for the analyst interface; briefings to analysts, scientists, and NGA leadership
- The interactive Stage 0 explorer below, built to make the screening construction inspectable
08Result
- The government's static screen became a configurable spatial universe with documented inclusion rules, so the same methodology serves different missions without altering evidence.
- Ownership became a calibrated probability with retained uncertainty and conflict, designed for empirical calibration in the pilot rather than asserted.
- Geometric exploitation potential entered as an independent analytical dimension, giving analysts a second axis that proximity alone could not supply.
- The specification, prototypes, and pilot form a reusable baseline: each stage can be inspected, recalibrated, replaced, or independently validated without redesigning the system.
- NGA's Director of Counterintelligence was briefed on the methodology; the Washington, DC validation pilot is in progress and will refine thresholds, parameters, and requirements from empirical behavior.
09Tradeoffs and limitations
- Independence assumptions. Noisy-OR and network structure assume partial independence among indicators; correlated sources can overstate confidence, which sensitivity analysis and conflict monitoring are designed to expose.
- Mass assignment. Dempster–Shafer requires source-reliability estimates to assign belief; those estimates are themselves calibrated in the pilot, not known in advance.
- Compute versus coverage. Monte Carlo geometry is expensive at inventory scale; the Stage 0 universe exists to spend that cost only where screening justifies it.
- Labeled outcomes are scarce. Calibration and false-negative analysis depend on adjudicated ground truth that accumulates slowly; early confidence tiers must be read as design targets.
- Adversarial obscuration. Deliberately layered ownership can defeat evidence collection; the methodology surfaces this as uncertainty rather than resolving it.
- Analyst workload. Every threshold creates review volume; the pilot measures that volume before thresholds are fixed.
- Illustration versus pilot. The explorer on this page runs on a synthetic terrain scene to make the construction legible; it is not the pilot and carries no pilot results.
10What this demonstrates
11Architecture and interactive Stage 0 explorer
Stage 0, computed live
The government's screen was a fixed statutory circle. Stage 0 replaces it with 𝒰 = 𝔅 ∪ 𝔓 ∪ 𝔖: the statutory buffer, a shadow buffer for positional uncertainty, and four physics-informed exploitation envelopes computed against terrain and the road network. At every azimuth the MTEBB envelope takes the farthest modality. Everything below is computed live from a synthetic terrain scene: a Monte Carlo viewshed, a knife-edge radio link budget, a road-network isochrone, and a UAS feasibility model that carries wind, endurance, command-link line of sight, and the geofence assumption. Click a parcel to trace exactly why it entered the universe.
The same four modalities computed on real terrain, with line of sight, viewshed, chokepoint ranking, and UAS approach exposure in three dimensions: open the Terrain Lab
How to read the scene
The cardinal point is the protected facility. The dashed circle is 𝔅, the dotted circle is 𝔖, and each coloured hull is one modality's reach. The thick cardinal line is the MTEBB envelope; the faint fill behind everything is 𝒰. Parcels are coloured by the pathway that admitted them; excluded parcels stay hollow. Toggling a pathway removes it from 𝒰 and from the envelope, so counts and area respond immediately.
Model notes
- VIS viewshed
- Radial line of sight repeated over Monte Carlo draws of target height, observer height, and correlated terrain error; admission by the probability of an unobstructed sight line.
- RF propagation
- Free-space path loss plus single knife-edge diffraction from the dominant obstruction of the first Fresnel zone, plus a clutter term, against a link budget.
- MOB network reach
- A Dijkstra isochrone over the terrain grid with road speed on the network and slope-penalised off-road speed elsewhere.
- UAS feasibility
- Round-trip range from usable endurance, cruise airspeed, and the wind vector resolved into along-track and crosswind components, plus observation standoff; piloted missions add command-link range and line of sight from the operator.