Creating a Repository
A Crab repository is a standard git repository with a cloud storage backend for large files. Crab handles the wiring between git and your bucket in a short setup flow.
Create the repository
mkdir my-project && cd my-project
crab configure s3://my-bucket/my-project
crab ship . -m "initial commit"crab configure selects the cloud provider, checks credential discovery,
creates the Git and Crab configuration, and detects large-file tracking rules.
Use the interactive form if you prefer prompts:
mkdir my-project && cd my-project
crab configureWhat You Need
Before creating a Crab repository, you need:
- A cloud storage bucket — S3, GCS, or Azure Blob Storage
- Credentials — Access to write to that bucket (AWS keys, GCP service account, or Azure credentials)
You do not need to run git init separately — crab configure handles that automatically.
The bucket/container itself must already exist. Crab creates the repository layout inside it, but it does not create cloud accounts, buckets, IAM roles, or long-lived credentials. After signing in to the provider, verify the complete setup with:
crab doctorGuided configuration
cd my-project
crab configure s3://my-bucket/my-projectThis command:
- Selects or infers S3, GCS, or Azure Blob Storage
- Finds credentials through the provider SDK chain; use
--aws-profile <name>to select an AWS profile - Runs
git initif no Git repository exists - Creates
.crab/andcrab.tomlwith remote configuration - Registers the filter and diff drivers in
.git/config - Scans for large files and writes matching
.gitattributesrules
Run crab configure <REMOTE> --dry-run to preview the setup plan. The lower-level
crab init and crab setup commands remain available when you want to perform
the two phases separately.
Example output
Initialized git repository in /home/user/my-project
Filter driver installed
Auto-tracked 3 extension(s) in .gitattributesWhat Crab Creates
| Created | Purpose |
|---|---|
.git/ | Git repository (created if missing) |
crab.toml | Committed remote, provider, tracking, hydration, prefetch, and workflow policy |
.crab/local.toml | Uncommitted machine settings such as an AWS profile selector and cache tuning |
.crab/staging/ | Uncommitted chunks awaiting push; may contain the only unpublished copy |
.crab/cache/ | Rebuildable local cache data |
.git/config changes | Registers the Crab filter driver |
.gitattributes | Tracking rules for large file extensions (auto-generated) |
Auto-Tracking
crab setup automatically detects files that should be managed by Crab:
- Files larger than 1 MiB trigger tracking for their extension
- Well-known binary formats (
.safetensors,.bin,.onnx,.parquet,.h5, etc.) are always tracked when found
This means you rarely need to manually run crab track — the common case is handled automatically. If you need to add more patterns later, use:
crab track '*.custom-extension'To skip auto-tracking (e.g., in CI or when you want full manual control):
crab setup --no-auto-trackURL Format
Crab URLs follow the pattern crab://<bucket>/<repo-path>:
crab://my-bucket/my-project
crab://company-data/team-ml/experiment-42
crab://us-west-2-storage/repos/frontend-assetsThe <repo-path> isolates this repository's data within the bucket. Multiple repositories can share a single bucket with different paths.
Available provider adapters
| Provider | URL Format | Example |
|---|---|---|
| AWS S3 | s3://<bucket>/<path> | crab configure s3://ml-data/models |
| Google Cloud Storage | gs://<bucket>/<path> | crab configure gs://ml-data/models |
| Azure Blob Storage | az://<container>/<path> or azure://<container>/<path> | crab configure azure://ml-data/models |
| S3-compatible (MinIO, R2, Ceph) | crab://<bucket>/<path> | crab configure crab://my-minio-bucket/data --provider s3 |
This table documents accepted configuration syntax. It does not mark those services as release-qualified.
How Provider Detection Works
For crab configure and crab init, Crab infers the provider from a provider-prefixed input URL
(s3://, gs:///gcs://, or az:///azure://) or uses the explicit
provider flag (--provider for configure, --storage-provider for init). Provider-prefixed inputs are normalized to a
canonical crab:// Git remote. For an unprefixed crab:// URL, set the
provider explicitly when it is not S3:
[auth]
storage_provider = "gcs" # "s3" | "gcs" | "azure" | "auto"The CRAB_STORAGE_PROVIDER environment variable is also accepted for
credential discovery. Its values are case-insensitive:
| Value | Provider |
|---|---|
s3 (default) | AWS S3 / S3-compatible |
gcs, gs, google | Google Cloud Storage |
azure, az, abs | Azure Blob Storage |
S3-Compatible Endpoints
For S3-compatible stores like MinIO, Cloudflare R2, or Ceph, select the s3
provider and set the service's S3 API endpoint and credentials before running
crab configure:
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=us-east-1
export AWS_ENDPOINT_URL=https://s3.example.com
export AWS_VIRTUAL_HOSTED_STYLE_REQUEST=false
crab configure crab://my-bucket/my-project --provider s3
crab doctorUse the region required by the service; Cloudflare R2 uses auto. For a trusted
plain-HTTP development endpoint, also set AWS_ALLOW_HTTP=true. Crab reads the
endpoint from AWS_ENDPOINT_URL, so credentials and endpoint URLs do not
belong in .crab.toml.
See Static Credentials for complete Cloudflare R2 and local MinIO/RustFS examples.
After Initialization
Once initialized, start adding files:
# Add your large files (auto-tracked extensions are already configured)
crab add .
# Commit and push
git commit -m "Initial commit with large files"
git pushOr use the one-shot shorthand:
crab ship . -m "Initial commit with large files"For Collaborators
When a teammate has already set up a Crab repository, you have two options:
Option A: crab clone (recommended)
crab clone crab://my-bucket/my-projectThis handles everything: git clone, filter driver, tracking rules, and optional hydration.
Option B: Global install + regular git clone
# One-time setup (works for all repos):
crab install --global
# Then clone normally:
git clone <url>With the global filter driver installed, any repo that has .gitattributes with filter=crab rules will work automatically.
Re-initialization
Running crab init on an already-initialized repository is safe and idempotent. It refreshes the filter driver configuration without overwriting your existing config or staging data. Run crab setup separately when you want to rescan for new large-file patterns.
Troubleshooting
"invalid URL" error — Ensure the URL follows crab://<bucket>/<path>. The bucket name must not contain slashes.
Bucket not found — Create the bucket/container or correct the bucket name in
the remote URL, then rerun crab configure.
Repository not initialized — The bucket exists, but the Crab layout does
not. Run crab configure <REMOTE> to create it.
Access denied or no credentials — Configure the selected provider's
credentials and grant the active identity the required bucket and
repository-prefix permissions. Run crab doctor to verify access before
retrying.
Filter driver not registered — Run crab doctor for a full health check. If the filter is missing, crab init or crab setup will re-register it.
Next Steps
- Track file patterns — Manually configure additional patterns
- Add and push files — Stage tracked files for push
- Clone a repository — Share with collaborators