← Work

python · playwright · sqlite

InfoPath Migration Pipeline

A resumable Python pipeline that moved 6M+ files from SharePoint 2007 to SharePoint Online, converting InfoPath forms to PDF.

Problem

A legacy SharePoint 2007 estate held millions of InfoPath XML forms that had to reach SharePoint Online as PDFs, with their metadata preserved. At that scale a run will be interrupted and some files will fail. The real question is whether you can say which ones, and pick up exactly where you stopped.

Constraints

  • 6M+ files, converted at high throughput.
  • Metadata extracted and preserved with each document.
  • Every file accounted for: done, waiting to retry, or failed for a stated reason.

Decisions

  • Three stages — index → schedule → process — so discovery, planning and conversion can each be rerun on their own.
  • Durable task and audit state in SQLite, written atomically. A restart continues from the recorded state instead of starting over.
  • An explicit error taxonomy: retryable errors go back in the queue; permanent errors are recorded and stop retrying.
  • Playwright renders the forms to PDF.

Verification

  • A golden-corpus test suite checks conversions against known-good output.
  • Worker-sizing benchmarks set concurrency from measurement, not guesswork.

Outcome

6M+ files migrated from SharePoint 2007 to SharePoint Online, with audit state behind every one of them.

Next: Enterprise Search Connector →