Migrating from DVC to Crab
This guide covers the repository-aware migration of a DVC project into
Crab's workflow format (crab.yaml). Migration is deliberately two things at
once: a pipeline conversion and a source-inventory protocol. It records DVC
metadata, pointers, cache objects, remotes, materialized data, lock records,
and run-cache state before it can publish Crab state. It never deletes DVC
state.
Table of Contents
- Overview
- Running the migration
- Conversion rules
- Unsupported features
- Manual steps after migration
- Example: 5-stage DVC pipeline
Overview
Crab covers the core DVC pipeline concepts, but it is not a general DVC replacement until the documented provider and data-command gates are met. Most ordinary pipeline fields map directly:
- Stages, deps, outs, params, metrics, plots — same semantics.
foreachandmatrix— same syntax.vars:and${...}templating — same syntax.wdir:andfrozen:— direct equivalents.
The migration tool first validates the converted YAML, then (unless
--stdout or --plan is used) inventories the project and writes a durable
journal under .crab/workflow/migration/. It transfers only sources that can
be verified, in this order: materialized output, local DVC cache, then a
qualified remote provider. A missing, corrupt, unknown, secret-bearing, or
unsupported record blocks cutover.
Running the migration
# From the repo root (where dvc.yaml lives):
crab migrate from-dvc
# From a different directory:
crab migrate from-dvc --dir path/to/project
# Print to stdout instead of writing crab.yaml:
crab migrate from-dvc --stdout
# Inventory without mutating YAML, Git, Crab data, or the journal:
crab migrate from-dvc --plan --json
# Resume a transfer only when the source inventory fingerprint is unchanged:
crab migrate from-dvc --resume
# Record an explicit, credential-free remote destination (still needs live proof):
crab migrate from-dvc --remote-map origin=crab://bucket/projectThe repository-aware command:
- Locates and parses
dvc.yaml. - Inventories recursive DVC metadata, standalone
.dvcpointers, cache roots and.dirmanifests, config precedence, remotes, materialized outputs, lock records, ignore files, and run-cache records. - Converts the pipeline and validates the generated
crab.yamlbefore any mutation. - Writes an atomic, resumable journal and verified Crab-owned source objects.
- Publishes canonical Crab pointers,
crab.lock, and the generated YAML only after source verification succeeds. - Reports every finding and the
safe_to_remove_dvcboolean. That boolean is true only after clean-clone restore evidence and all cutover gates pass.
--plan performs only steps 1–3 and emits no file, Git index, Crab data, or
journal mutation. --stdout is conversion-only and cannot report a safe
cutover. --remote-map records an intended destination; it does not claim
that the destination was populated or live-verified. Remote mappings remain
blocking until that evidence exists.
After migration, validate the result:
crab run --validateConversion rules
| DVC field | Crab equivalent | Notes |
|---|---|---|
cmd: (string) | cmd: (string) | Direct copy |
cmd: (list) | cmd: (list) | Preserved; each entry runs in a fresh native shell |
deps: | deps: | Direct copy |
outs: | outs: | Subfields mapped (see below) |
params: | params: | Direct copy |
metrics: | metrics: | Direct copy |
plots: | plots: | Simplified (no rendering config) |
wdir: | wdir: | Direct copy |
frozen: | frozen: | Direct copy |
always_changed: | nondeterministic: | Renamed |
foreach: | foreach: | Syntax match |
matrix: | matrix: | Syntax match |
vars: | vars: | Direct copy |
${...} | ${...} | Same syntax |
desc: | desc: | Direct copy |
meta: | meta: | Direct copy |
Output subfields
| DVC output field | Crab equivalent |
|---|---|
cache: true/false | cache: true/false |
persist: true | persist: true |
push: false | push: false |
checkpoint: true | Rejected |
remote: <name> | Blocked until mapped and verified |
Command list conversion
DVC allows cmd: as a list of strings. Crab preserves that list and executes
each entry in order in a fresh shell. Shell syntax is not translated between
operating systems: Unix uses /bin/sh -c, while Windows uses
cmd.exe /D /S /C. Use Crab's argv form for a portable command.
# DVC
cmd:
- mkdir -p output
- python build.py
- python validate.py
# Crab (after migration)
cmd:
- mkdir -p output
- python build.py
- python validate.pyalways_changed → nondeterministic
DVC's always_changed: true becomes nondeterministic: true in crab.
Same semantics: the stage always re-executes regardless of input hashes.
Unsupported features
These boundaries are not warnings that can be ignored. A semantic or source
finding that would make the converted project unsafe is fatal and leaves the
previous tracked state intact. Informational findings may be reported, but
they do not make safe_to_remove_dvc true.
| DVC feature | Workaround |
|---|---|
push: false on outputs | Preserved in the converted declaration; remote stage-cache publication remains disabled for that stage. |
remote: per output | The redacted remote is inventoried. Supply an explicit --remote-map; destination population and live verification are still required before cutover. |
| DVC checkpoints | Rejected with dvc_checkpoint_unsupported; Crab checkpoints are explicit experiment lineage and are not rewritten to persist. |
artifacts: declarations | Preserved and validated as catalog metadata. A configured primary Crab remote provides immutable publication and CAS promotion; migration remains unsafe until clean-clone and artifact-GC reachability gates pass. |
live: (DVCLive integration) | Use crab's metrics and plots directly. DVCLive callbacks write standard JSON/CSV that crab can track. |
| DVC import provenance or unsupported providers | The inventory preserves only credential-free provenance. SSH/SFTP, HDFS/WebHDFS, WebDAV, Drive, and OSS remain unsupported until their live provider gates pass; unsupported provenance blocks cutover. |
Hydra composition (dvc exp run with Hydra) | Use crab exp run --set key=value for param overrides. For complex Hydra configs, run Hydra as part of your stage command. |
Manual steps after migration
After running crab migrate from-dvc:
-
Review findings and the journal. Run
crab migrate from-dvc --plan --jsonfirst, then inspect.crab/workflow/migration/dvc.json. Every blocking finding must be resolved; a mapped remote is not resolved until its destination has been populated and live-verified. -
Validate the pipeline.
crab run --validateFix any schema errors or undefined template references.
-
Verify the cutover report. A repository-aware migration recomputes Crab hashes and writes
crab.lockfrom verified bytes; it never copies DVC MD5 or ETag values into Crab identities. If migration is blocked, do not hand-write a lockfile to bypass the journal. Resume withcrab migrate from-dvc --resumeafter correcting the source or mapping. -
Update CI scripts. Replace
dvc reprowithcrab runanddvc pushwithcrab run --cache-push. -
Update
.gitignore. DVC adds entries like/model.pklfor tracked outputs. Crab uses the same pattern — your existing.gitignorelikely works as-is. -
Keep DVC artifacts. This command never deletes DVC state. Keep
dvc.yaml,dvc.lock,.dvc/, cache objects, remotes, and pointer files until the report explicitly sayssafe_to_remove_dvc: trueand includes a clean-clone restore with byte-, tree-, and mode-identical verification. Crab has no delete flag; any later removal is a separate manual change. -
Commit the migration.
git add crab.yaml crab.lock .gitignore git commit -m "Migrate pipeline from DVC to crab"
Example: 5-stage DVC pipeline
Before (dvc.yaml)
vars:
- codedir: src
- datadir: data
stages:
download:
cmd: "python ${codedir}/download.py --out ${datadir}/raw.csv"
deps:
- ${codedir}/download.py
outs:
- ${datadir}/raw.csv
always_changed: true
clean:
cmd: "python ${codedir}/clean.py"
deps:
- ${codedir}/clean.py
- ${datadir}/raw.csv
outs:
- ${datadir}/clean.parquet
featurize:
cmd: "python ${codedir}/featurize.py"
deps:
- ${codedir}/featurize.py
- ${datadir}/clean.parquet
params:
- features.window_size
- features.columns
outs:
- ${datadir}/features.parquet
train:
cmd:
- mkdir -p models
- python ${codedir}/train.py
deps:
- ${codedir}/train.py
- ${datadir}/features.parquet
params:
- model.lr
- model.epochs
- model.arch
outs:
- models/model.pkl:
persist: true
- models/checkpoints/:
push: false
metrics:
- metrics/train.json
evaluate:
cmd: "python ${codedir}/evaluate.py"
deps:
- ${codedir}/evaluate.py
- models/model.pkl
- ${datadir}/features.parquet
metrics:
- metrics/eval.json
plots:
- metrics/roc.csvAfter (crab.yaml — generated by crab migrate from-dvc)
vars:
- codedir: src
- datadir: data
stages:
download:
cmd: "python ${codedir}/download.py --out ${datadir}/raw.csv"
deps:
- ${codedir}/download.py
outs:
- ${datadir}/raw.csv
nondeterministic: true
clean:
cmd: "python ${codedir}/clean.py"
deps:
- ${codedir}/clean.py
- ${datadir}/raw.csv
outs:
- ${datadir}/clean.parquet
featurize:
cmd: "python ${codedir}/featurize.py"
deps:
- ${codedir}/featurize.py
- ${datadir}/clean.parquet
params:
- features.window_size
- features.columns
outs:
- ${datadir}/features.parquet
train:
cmd: "mkdir -p models && python ${codedir}/train.py"
deps:
- ${codedir}/train.py
- ${datadir}/features.parquet
params:
- model.lr
- model.epochs
- model.arch
outs:
- models/model.pkl:
persist: true
- models/checkpoints/:
push: false
metrics:
- metrics/train.json
evaluate:
cmd: "python ${codedir}/evaluate.py"
deps:
- ${codedir}/evaluate.py
- models/model.pkl
- ${datadir}/features.parquet
metrics:
- metrics/eval.json
plots:
- metrics/roc.csvMigration report
Migration Report
==================================================
Stages converted: 5
Output written to: crab.yaml
Warnings: none
safe_to_remove_dvc: false
Blocking findings:
dvc_remote_clean_clone_unverified
==================================================DVC checkpoint outputs are rejected until they can be represented by Crab
experiment lineage; they are never downgraded to persist or silently treated
as ordinary cached outputs. Ordinary push: false output settings are
preserved by the converter.