Web data built for back-testing, not just dashboards.
Investment research needs more than a current snapshot. We build web-derived datasets on prices, listings, hiring, footprints and availability with point-in-time history, documented backfill and controls for survivorship bias, so signals can be tested honestly before they inform a position.
Four steps from sources to signal.
Scope the signal and universe
We work with your analysts to define the companies, sources and metrics, such as price changes, job postings, store counts or SKU availability, and the tickers they map to.
Capture point-in-time
Every record is stored with its capture timestamp and never overwritten. Corrections are appended as new versions, so you can reconstruct exactly what was knowable on any date.
Backfill with documentation
Where history is needed, we recover it from archived public pages and historical listings. Backfilled records are flagged separately from live captures, with method and coverage documented.
Control for survivorship and coverage
Delisted products, closed locations and inactive companies stay in the panel with end dates. Panel size and source changes are logged so coverage drift can be adjusted for.
What we capture.
Every field is validated, normalised and documented in a data dictionary you can share with your analysts.
Capture timestamp
When each observation was recorded, the basis of point-in-time queries.
Record version
Version history for any corrected or updated value.
Backfill flag
Whether a record was captured live or reconstructed from archives.
Entity and ticker mapping
Company, brand and ticker identifiers for each record.
Active and end dates
When a product, store or listing entered and left the panel.
Panel coverage metrics
Count of tracked entities and sources per period.
What you receive.
- Point-in-time dataset delivered to your warehouse or cloud storage
- Backfill methodology and coverage note
- Entity-to-ticker mapping table, maintained over time
- Panel coverage and source change log
- Weekly or daily incremental feed with a stable schema
An illustrative extract.
| Ticker (placeholder) | Region | As of date | Active stores | Opened | Closed | Source |
|---|---|---|---|---|---|---|
| TICK-A | United States | 2026-03-31 | 1,284 | 12 | 5 | Live capture |
| TICK-A | Canada | 2026-03-31 | 146 | 2 | 0 | Live capture |
| TICK-B | United Kingdom | 2025-12-31 | 412 | 3 | 9 | Backfill |
| TICK-B | United Kingdom | 2026-03-31 | 408 | 4 | 8 | Live capture |
| TICK-C | Australia | 2026-03-31 | 231 | 6 | 1 | Live capture |
Illustrative rows. Sources, markets and fields are agreed with you during scoping.
More in market intelligence
Relevant industries
Underlying services
What does point-in-time mean in practice?
Every value is stored with the time it was captured, and later corrections are added as new versions rather than overwriting. You can query the dataset as it stood on any past date, which helps prevent look-ahead bias in back-tests.
How do you handle survivorship bias?
Entities that disappear, such as delisted products, closed stores or removed job postings, remain in the history with an end date. The panel therefore reflects the full universe at each point, not only the survivors.
Can you share a sample for evaluation?
Yes. We provide a free sample in 24–48 hours, and can sign an NDA on request. For investment use we include the methodology note so your team can assess backfill and coverage before committing.
