Data On Demand, one-stop data solution
Markets Local sources. Your time zone.

Teams in five regions rely on us for local-language, multi-currency data.

Global coverage
Company A decade of data expertise.

An in-house team of 20+ engineers and analysts serving clients in 12+ countries.

About us
Core service

Web corpora curated for training, not just collected.

Models are only as good as the data behind them. We build domain and language-specific corpora from public web sources, with provenance recorded, licence signals respected, duplicates removed and quality filters applied, ready for labelling or fine-tuning.

The challenge

Generic crawl dumps are noisy, repetitive and heavy with boilerplate, spam and harmful content. Teams building models for Arabic, regional commerce or specialist domains struggle to find enough clean, well-documented text and structured examples.

With Data On Demand

You receive a corpus shaped around your model's purpose, with each document traceable to its source and licence signals. Quality and safety filtering is documented, so your team can explain what went into training and why.

What’s included

Everything needed to run it in production.

01

Source curation

We select domains and sources by relevance, language and quality, and record licence and robots signals for each one.

02

Licence-aware collection

Content with restrictive licence terms or opt-out signals is excluded or flagged, according to the policy you set.

03

Near-duplicate removal

Exact and near-duplicate documents are removed with hashing and similarity methods, reducing memorisation and wasted compute.

04

Quality and toxicity filtering

Boilerplate, spam, low-information pages and toxic content are filtered using configurable classifiers and thresholds.

05

Personal data minimisation

Emails, phone numbers and other identifiers are detected and masked or removed before delivery.

06

Labelling-ready formats

Output in JSONL, Parquet or instruction-pair formats with metadata fields your labelling tools and training pipelines can ingest directly.

Sample output

What lands in your systems.

Typical fields
doc_idsource_domainlanguagelicence_signaltoken_countquality_scoretoxicity_scorededupe_cluster_id
Corpus manifest: documents after de-duplication and filtering
doc_idsource_domainlanguagelicence_signaltokensquality_scoretoxicity_scorestatus
d_0019a4news-site-a.examplearopen, attribution1,2840.910.02kept
d_0019b7gov-portal-b.exampleenpublic sector licence3,9060.880.01kept
d_0019c2forum-c.examplemsunspecified4120.370.04dropped_low_quality
d_0019d0retail-d.examplept-BRterms reviewed6550.790.01dropped_near_duplicate
d_0019e8blog-e.examplees-MXopt-out signal1,1200.840.03excluded_licence

Illustrative rows. Your schema, field names and formats are agreed during scoping.

Use cases by team

Who uses it, and for what.

Machine learning teams fine-tune language models on a clean Arabic or multilingual corpus. Documented filtering makes evaluation results easier to interpret.

How it runs
  1. Scope. Tell us the sources, fields and frequency. We confirm feasibility within a day.
  2. Free sample. A real sample from your own target source, in your format.
  3. Build. Engineers build extractors tuned to each source. No generic templates.
  4. Validate. Automated and manual QA on every run before anything ships.
  5. Deliver and monitor. Scheduled delivery, monitored pipelines, fast fixes when sites change.
How we work
FAQ

Questions about ai training data

Can’t find your answer? Ask an engineer

How do you handle licensing and opt-outs?

We record licence and robots signals for each source and apply the policy you choose, such as excluding content with opt-out signals or restrictive terms. The manifest lists every source and its status for your legal review.

How do you remove duplicates?

We remove exact duplicates by hashing and near-duplicates using similarity clustering at document and paragraph level. Each kept document carries a cluster ID so you can inspect what was removed.

Can you filter toxic or harmful content?

Yes. Configurable classifiers score each document for toxicity and quality, and you set the thresholds. Filtered samples are available for review so you can tune the settings.

What formats do you deliver?

Typically JSONL or Parquet with text and metadata fields, or instruction-response pairs for fine-tuning. We can match the schema your labelling or training pipeline expects.

Start with proof

See your own data before you commit.

Name the sources and fields you need. Within 24–48 hours you receive a real sample from your target sites, in your format, free of charge.

Request a free sample Talk to a data engineer Sample in 24–48 hours · NDA on request · Any format, any schedule