Skip to content

Services

Data migration and publishing

We move data out of legacy systems, spreadsheets and archives into modern databases and formats, check that nothing was lost on the way, and publish it when others need to use it.

The problem

Valuable data often sits where only one system or one person can reach it: an ERP that is being retired, a database nobody wants to touch, a folder of spreadsheets, a file archive, a script on a former employee’s laptop. Copying it is the easy part. The hard part is keeping the meaning of each field, proving that every record arrived intact, recording where the data came from, and making sure it stays current once the project ends.

What we deliver

  • A mapping from the old structure to the new one, agreed with the people who use the data
  • Data loaded into the destination you choose, whether that is a new application, a cloud database or warehouse, or open formats such as Parquet and Iceberg
  • Validation that reconciles record counts and key totals between the old system and the new one
  • A data dictionary and a record of where each dataset came from
  • A cutover plan, or an update pipeline when the source keeps changing
  • An API, catalog or download site when the data needs to be shared
  • Handoff documentation, or a maintenance retainer if you would rather we keep it running

Public data is part of this work too. We rebuilt FEMA’s Future Risk Index after it was taken offline and engineered HIFLD Next, the public catalog of more than 400 preserved federal infrastructure datasets.

Where it applies

Companies changing systems
Records moved from an old ERP, CRM or homegrown database into the new one, with every table reconciled before cutover.
Teams outgrowing spreadsheets
Shared spreadsheets and Access files consolidated into a database your team and your reporting tools can rely on.
Government
Legacy databases, file archives and GIS layers converted to open formats such as GeoParquet, Cloud-Optimized GeoTIFF and Zarr, with STAC metadata and a pipeline that keeps them updated.
Research labs
Analysis scripts and manual handoffs moved into reproducible pipelines, and data prepared for repository deposit.
Open-data organizations
Public datasets republished and maintained after the original source goes offline.
Companies built on public data
Versioned copies of the public datasets your product depends on, so a change or takedown upstream does not break it.

Related work

Further reading

Start with one dataset, one report or one document type.

We will look at the systems and records involved and recommend a first step.