Playbook

Schema change detection: stop a renamed column from breaking your model silently.

A portal renames a column. The lookup returns blanks, the sheet still looks fine and the model runs on it for a month.

The problem

Public data portals change layouts, column names and formats without notice. Copy-paste workflows and simple scripts keep running, but the numbers they produce are wrong or empty, and nobody notices until a decision has been made on them.

The playbook

  1. Define the expected shape of every source: columns, types and roughly how many rows.
  2. Check before storing: compare each run’s shape with the expected one and stop on a mismatch.
  3. Report, do not swallow: a layout change becomes an alert with what changed, not a blank cell.
  4. Normalise explicitly: convert “12.4%”, “(3.2)” and “--” to 12.4, −3.2 and blank with written rules.
  5. Match by identifiers, such as registration number or ISIN, not names.
  6. Reconcile counts against totals shown at the source.
  7. Version every run and keep the previous value of anything that changed.
  8. Keep lineage on every record: source, fetch time, run and version.

Checks worth automating

CheckCatches
Column names and orderRenamed or moved columns
Row count against last runTruncated or partial pulls
Share of blanks per columnLookups that silently failed
Totals against the sourceDuplicates and missing rows

How to measure it

Runs stopped by a check, time from a source change to an alert, and errors found after data was used. The target for the last one is zero.

How Terminal X does it

Terminal X runs data agents on public market and regulatory sources for portfolio managers, brokers and research desks. Four standard agents cover the SEBI Portfolio Manager Monthly Report for every registered portfolio manager, PMS strategy performance from APMI against Nifty 50 or Nifty 500, NSE ETF liquidity, and Nifty strategy-index constituents. Each agent works the portal in a real browser with parallel windows, extracts every row and sub-report, turns text like “12.4%”, “(3.2)” and “--” into numbers or blanks, matches names by registration number or ISIN, reconciles counts, marks empty filings as “no data” rather than zero and compares each run with the last. Every record keeps its source, fetch time, run and previous version; layout changes at the source are reported, not swallowed. Results arrive as the same Excel export every time, a REST API and alerts on new data, and other sources can be added as custom agents.

This page describes data collection and research workflow only. It is not investment advice.

Questions

What is schema change detection?

Checking each data pull against the expected columns, types and row counts and raising an alert when the source layout changes, instead of storing wrong or empty values.

How do I stop silent failures in market data?

Validate shape and counts before storing, normalise values with explicit rules, match by identifiers, version every run and keep lineage on every record.

See it on your own data. Terminal X — Filings and market data, collected. Book a 30-minute working session with an engineer.

General guidance, current as of the date above. Figures and examples are illustrative unless a source is linked.