Project Configuration (crab.toml)
The crab.toml file lives in your repository root and declares how Crab should behave for this project. It's committed to git so every collaborator inherits the same configuration.
Full Schema
version = 1
# Required: cloud storage location for xorbs and manifests
[remote]
url = "crab://my-bucket/my-repo"
# Optional: file patterns to track with Crab
[track]
patterns = ["*.bin", "*.safetensors", "*.parquet", "datasets/**"]
# Optional: hydration behavior on clone/checkout
[hydrate]
default = "lazy" # "lazy" | "eager"
auto_patterns = ["*.py", "*.rs", "*.toml", "README*", "LICENSE*"]
# Optional: mirror mode (GitHub + Crab coexistence)
[mirror]
origin_remote = "origin"
crab_remote = "crab"
# Optional: credential hints
[auth]
storage_provider = "s3" # "s3" | "gcs" | "azure" | "auto"
# Optional: shared hydration sets
[prefetch.profiles.always]
paths = ["README.md", "src/**"]
# Optional: shared workflow policy
[workflow]
enabled = true
discover = "root"
parallelism = 4
Sections
[remote] (required)
The only required section. Specifies where Crab stores chunked file data.
[remote]
url = "crab://my-bucket/my-repo"Supported URL schemes:
crab://— AWS S3 (or S3-compatible)gs://— Google Cloud Storageaz://— Azure Blob Storage
[track] (optional)
Declares which file patterns Crab manages. These are synced to .gitattributes with the filter=crab attribute.
[track]
patterns = ["*.bin", "*.safetensors", "*.onnx", "datasets/**"]If omitted, crab setup auto-detects large files (>1 MiB) and well-known binary extensions.
[hydrate] (optional)
Controls what happens after clone or checkout.
[hydrate]
default = "lazy"
auto_patterns = ["*.py", "*.rs", "*.toml", "README*"]| Field | Values | Default | Description |
|---|---|---|---|
default | "lazy", "eager" | "lazy" | Whether to hydrate all files on clone |
auto_patterns | Array of globs | [] | Always hydrate these patterns regardless of default |
With default = "lazy", files remain as pointers until explicitly hydrated. With default = "eager", all tracked files are hydrated immediately after clone.
[mirror] (optional)
Enables mirror mode for GitHub/GitLab + Crab coexistence. See Mirror Mode for the full guide.
[mirror]
origin_remote = "origin"
crab_remote = "crab"When present, crab init (re-apply mode) installs pre-push and post-checkout hooks automatically.
[auth] (optional)
Explicit credential hints. Rarely needed — Crab's credential discovery chain finds credentials automatically from environment variables, cloud SDK configs, and instance metadata.
[auth]
storage_provider = "s3"AWS profile names are machine-specific and do not belong in committed project
configuration. Select one with crab configure --aws-profile ml-team, crab config set auth.aws_profile ml-team, or AWS_PROFILE=ml-team.
Precedence
Configuration resolves in this order (highest priority first):
- CLI flags —
--pattern,--eager,--mirror, etc. crab.toml— project-level defaults- Built-in defaults — lazy hydration, auto-detection
crab.toml vs .crab/local.toml
crab.toml | .crab/local.toml | |
|---|---|---|
| Location | Repo root | .crab/ directory |
| Purpose | Shared project policy | Machine-local settings |
| Committed to git | Yes | No |
| Edited by | Users, crab init | Crab automatically |
Think of crab.toml as "what this repo needs" and .crab/local.toml as
"how this machine accesses and runs it." Staging, caches, journals, and locks
also live under .crab/; staging may contain unpublished data and should not
be treated as a disposable cache.
The retired .crab.toml filename is not read. Rename it to crab.toml and
commit the rename.
Examples
ML Repository
[remote]
url = "crab://ml-artifacts/bert-finetune"
[track]
patterns = ["*.safetensors", "*.bin", "*.onnx", "*.pt", "datasets/**"]
[hydrate]
default = "lazy"
auto_patterns = ["*.py", "*.yaml", "requirements.txt", "README*"]Collaborators clone Git metadata and pointers, then hydrate the specific model checkpoint they need.
Monorepo with Large Assets
[remote]
url = "crab://company-assets/monorepo"
[track]
patterns = [
"assets/**/*.psd",
"assets/**/*.fbx",
"assets/**/*.blend",
"builds/**"
]
[hydrate]
default = "lazy"
auto_patterns = ["*.ts", "*.tsx", "*.json", "*.md", "*.css"]Developers get code hydrated immediately. Designers hydrate asset files on demand.
Mirror Mode (GitHub + Crab)
[remote]
url = "crab://team-bucket/our-project"
[track]
patterns = ["*.bin", "*.parquet", "models/**"]
[mirror]
origin_remote = "origin"
crab_remote = "crab"
[hydrate]
default = "lazy"
auto_patterns = ["*.py", "*.rs", "*.toml"]
[auth]
storage_provider = "s3"Code goes to GitHub via origin. Large files go to S3 via crab. The team's PR workflow stays unchanged.
How It's Generated
crab init <url> creates crab.toml automatically with:
[remote]from the provided URL[auth]when a storage provider is selected[mirror]if--mirrorflag was used
Run crab setup to add or refresh [track] patterns and the matching
.gitattributes entries.
Related
- Mirror Mode — GitHub + Crab coexistence
crab init— generatescrab.tomlcrab track— updates tracked file patterns