crab init
Initialize a new Crab repository.
Synopsis
crab init [OPTIONS] [URL]Description
crab init creates or validates the canonical v1 remote layout and empty
generation-0 manifest, then connects a local Git repository to that remote. It
creates or updates .crab/ and crab.toml, registers the Crab filter and
diff drivers, and adds a Git remote. It does not scan the working tree.
If no git repository exists in the current directory, crab init automatically
runs git init first — no need to initialize git separately.
Run crab setup afterward to scan for large files and write .gitattributes
tracking rules. With no URL, crab init re-applies the configuration from an
existing crab.toml.
For conceptual background, see Creating a Repository.
Arguments
| Argument | Required | Description |
|---|---|---|
[URL] | No | Cloud storage URL; omit it to re-apply an existing crab.toml |
Options
| Option | Default | Description |
|---|---|---|
--storage-provider <PROVIDER> | auto | Storage backend: s3, gcs, azure, or auto |
--gc-list-profile <PROFILE> | adaptive | Local bucket-GC policy: adaptive, cost, or latency |
--mirror <REMOTE> | — | Configure mirror mode using an existing Git remote |
--json | false | Emit structured JSON output |
--jsonl | false | Emit streaming JSONL output |
--log-level | — | Set log verbosity (error, warn, info, debug, trace) |
What It Does
- Creates a git repository (if
.gitdoesn't exist) - Creates
.crab/local.tomlfor machine-specific settings - Registers the filter and diff drivers in
.git/config - Adds or updates a Git remote for the canonical
crab://URL - Creates or validates the canonical v1 layout descriptor remotely
- Creates the generation-0 manifest atomically, or adopts the existing canonical manifest on repeat or concurrent initialization
There is no local-only selector. A non-empty prefix without a descriptor, a non-v1 descriptor, or a missing manifest during normal repository use fails closed. Reset isolated development data explicitly; push and clone do not convert or synthesize repository state.
Examples
New project from scratch
mkdir my-ml-project && cd my-ml-project
crab init crab://my-bucket/ml-models
# ✓ Initialized git repository
# ✓ Created .crab/local.toml
# ✓ Registered filter.crab driver
# Next: crab setupExisting git repo
cd my-existing-repo
crab init s3://team-bucket/my-repo
# ✓ Created .crab/local.toml
# ✓ Registered filter.crab driver
# Next: crab setupConfigure tracking separately
crab init crab://bucket/repo
crab setup --no-auto-track
# Install the filter driver without scanning for patternsChoose bucket-GC cost or latency policy
crab init --gc-list-profile adaptive crab://bucket/repoadaptive keeps small listings on one low-cost stream and switches large
namespaces to concurrent hash partitions. cost always minimizes LIST
streams; latency immediately uses partition parallelism. The setting is
local operational policy and does not change shared object keys.
Re-apply a committed project configuration
crab initAuto-Tracked Extensions
When auto-tracking is enabled (the default), crab setup tracks extensions that meet either criterion:
- Size threshold: Any file above 1 MiB triggers tracking for its extension
- Well-known binary formats: These extensions are always tracked when found, regardless of size:
| Domain | Extensions |
|---|---|
| ML/AI | .safetensors, .bin, .onnx, .pt, .pth, .h5, .hdf5, .pkl |
| Data | .parquet, .arrow, .feather, .npy, .npz, .zarr |
| Media | .fbx, .blend, .psd, .tiff, .exr, .dpx, .mov, .mp4, .wav |
| Archives | .tar, .gz, .zip, .zst, .lz4 |
| Databases | .db, .sqlite, .sqlite3 |
URL Format
Crab URLs follow the pattern crab://<bucket>/<repo-path>:
crab://my-bucket/my-project
crab://company-data/team-ml/experiment-42
s3://us-west-2-storage/repos/frontend-assetsRelated Commands
crab clone— clone an existing repositorycrab track— manually configure file patternscrab install— install filter driver globallycrab ship— one-shot add + commit + push