Data migration and publishing
We move data out of legacy systems, spreadsheets and archives into modern databases and formats, check that nothing was lost on the way, and publish it when others need to use it.
The problem
Valuable data often sits where only one system or one person can reach it: an ERP that is being retired, a database nobody wants to touch, a folder of spreadsheets, a file archive, a script on a former employee’s laptop. Copying it is the easy part. The hard part is keeping the meaning of each field, proving that every record arrived intact, recording where the data came from, and making sure it stays current once the project ends.
What we deliver
- A mapping from the old structure to the new one, agreed with the people who use the data
- Data loaded into the destination you choose, whether that is a new application, a cloud database or warehouse, or open formats such as Parquet and Iceberg
- Validation that reconciles record counts and key totals between the old system and the new one
- A data dictionary and a record of where each dataset came from
- A cutover plan, or an update pipeline when the source keeps changing
- An API, catalog or download site when the data needs to be shared
- Handoff documentation, or a maintenance retainer if you would rather we keep it running
Public data is part of this work too. We rebuilt FEMA’s Future Risk Index after it was taken offline and engineered HIFLD Next, the public catalog of more than 400 preserved federal infrastructure datasets.
Where it applies
- Companies changing systems
- Records moved from an old ERP, CRM or homegrown database into the new one, with every table reconciled before cutover.
- Teams outgrowing spreadsheets
- Shared spreadsheets and Access files consolidated into a database your team and your reporting tools can rely on.
- Government
- Legacy databases, file archives and GIS layers converted to open formats such as GeoParquet, Cloud-Optimized GeoTIFF and Zarr, with STAC metadata and a pipeline that keeps them updated.
- Research labs
- Analysis scripts and manual handoffs moved into reproducible pipelines, and data prepared for repository deposit.
- Open-data organizations
- Public datasets republished and maintained after the original source goes offline.
- Companies built on public data
- Versioned copies of the public datasets your product depends on, so a change or takedown upstream does not break it.
Related work
Further reading
- Migrating off a legacy system without losing history
- HIFLD Next: Restoring America’s Infrastructure Datasets
- Beyond Preservation: Stewarding America’s Critical HIFLD Infrastructure Data
- How We Prepared HIFLD for the Singularity
- Trump’s ‘climate’ purge deleted a new extreme weather risk tool. We recreated it