Need exact lineage?
Commit pointer blobs beside code so a Git ref identifies the complete project state.
Crab keeps Git as the timeline and your object store as the data plane. Version the whole project, move only the content a task needs, and keep every artifact tied to a commit.
commit a18f7e2
experiment/vision-v4
train.py
ordinary Git blob
model.safetensors
Crab pointer blob
chunk · deduplicate
verify · reconstruct
Find your fit
Crab uses one storage model across every team: pointers in Git, content-addressed chunks in object storage, and explicit materialization in the workspace.
Commit pointer blobs beside code so a Git ref identifies the complete project state.
Content-defined chunks let new versions reuse byte sequences already stored in the repository.
Clone lazily, hydrate selected paths, then dehydrate clean files when the task is done.
Use direct-storage mode with the object store and credential chain your team already operates.
Put weights, adapters, training data, configs, and metrics on one Git timeline without turning large payloads into ordinary Git blobs.
Where teams get stuck
Model artifacts often live under bucket names or registry tags that drift away from the training commit. Repeated checkpoints can also contain long byte ranges that earlier versions already uploaded.
The Crab workflow
Crab commits small pointer blobs to Git and stages the original bytes as content-defined chunks. On push, content-addressed deduplication can reuse chunks already present. A teammate can clone the history first, then hydrate only the model family needed for evaluation.
One model, three refs
Shared chunks stay content-addressed
main
base checkpoint
exp/vision
fine-tune
release/v4
promoted model
crab hydrate --profile=evalVersion notebooks and their inputs together so a branch, tag, or release identifies the complete analytical state—not just the code around it.
Where teams get stuck
A notebook may be committed while its CSV, Parquet, imagery, or simulation output is replaced in place. The code is reproducible; the input path is not.
The Crab workflow
Track selected data paths with Crab and commit their pointers beside the notebook. Each Git ref resolves to reconstruction metadata for the exact bytes. Researchers can inspect state with Crab, switch refs with Git, and hydrate only the paths required for the next analysis.
release/q3-analysis
one ref · complete analytical state
analysis.ipynb
Git blob
params.yaml
Git blob
train.parquet
Crab pointer
imagery/
Crab pointers
datasets/validation/hydrated
datasets/training/pointer
Switching refs changes both the analysis and the exact data snapshot it names.
crab hydrate 'datasets/validation/**'Keep textures, audio, video, CAD, and render outputs in the project history while each workstation materializes only its active set.
Where teams get stuck
Creative repositories grow faster than workstation disks. A fresh collaborator may need a few scenes or asset families—not every historical binary before the first edit.
The Crab workflow
A lazy Crab clone checks out pointers instead of every managed payload. Artists can hydrate by path, use a named profile, or mount a repository for on-demand reads when a supported NFS or FUSE backend is available. Clean files can be dehydrated back to pointers to reclaim disk space.
world-building / desert-city
Only the active scene is materialized
textures/
local
cinematics/
pointer
meshes/
local
audio/
pointer
references/
pointer
lighting/
local
Full project
object storage
Active set
local disk
Crab will not dehydrate a dirty file or replace it with a stale pointer.
crab dehydrate --allKeep large fixtures and release inputs on the same ref as the pipeline, then hydrate a deterministic manifest instead of pulling the entire repository payload.
Where teams get stuck
CI jobs often download a large shared fixture set even when a test shard touches only a handful of files. Ephemeral runners repeat that transfer unless the pipeline describes its working set.
The Crab workflow
Crab supports named prefetch profiles and newline-delimited hydrate manifests. Jobs can pre-warm the cache without changing the checkout, materialize their paths, and use JSON or JSONL output for stable automation and diagnostics.
lazy clone
history + pointers
hydrate
manifest paths
test
real fixtures
01tests/fixtures/auth.bin
02models/tiny.safetensors
03snapshots/linux/**
Warm before checkout
crab fetch can pre-load selected content without changing working-tree files.
crab hydrate --manifest .crab/manifests/ci.txtFor direct-storage repositories, developers and runners connect to your S3, GCS, Azure Blob, or S3-compatible bucket—without a Crab data server in the path.
Where teams get stuck
A new large-file service can introduce another data plane, credential model, scaling surface, and vendor boundary for the platform team to own.
The Crab workflow
Crab discovers provider-native credentials such as AWS profiles and roles, Google Application Default Credentials, or Azure managed identities and SAS credentials. Git objects, refs, chunks, and reconstruction metadata live under the configured repository prefix in your bucket.
Developer
crab binary
CI runner
crab binary
Object storage
Git objects · refs · xorbs · shards · indexes
crab doctorConfigure a repository against your bucket, track the paths that matter, and ship code and content together.