Skip to main content

Managing Repos

Once your Arraylake organization has been fully configured and you have installed the client library, you're ready to start managing data! 🎉

Create a Repo

For the purposes of this example, our org name will be earthmover. If running these commands interactively, replace earthmover with your org name.

For this example, we are going to create a Repo called ocean to hold oceanography data 🌊.

Navigate to your org's home page (for the org called earthmover this is https://app.earthmover.io/earthmover), then click New Repository.

Button to create a new repository in the web app.

The new repository button.

Type in the desired Repo name and hit Create.

Options when creating a new repository in the webapp.

The create repository form.

You can optionally add a repo description and/or metadata to your repo. Descriptions are limited to 255 characters, while repo metadata can be at most 4kB total. Metadata must be a mapping of key-value pairs where values are strings, numbers, lists, booleans, or None.

arraylake repo create earthmover/ocean \
--description "This is some oceanography data!" \
--metadata '{"type": ["climate", "coastal",
"environmental"], "model": "FVCOM"}'

You can also modify the description and/or metadata of an existing repo.

arraylake repo modify earthmover/ocean --description "This is an updated description for some oceanography data!" -a '{"source": "NOAA"}' -r "type" -u '{"model": "ROMS"}'

Where is Repo data stored?

Arraylake lets you configure the storage location for a Repo's data using org-level bucket configurations. Choose a specific bucket by providing bucket_config_nickname to create_repo. If not specified, the organization's default bucket is used.

Within the bucket and prefix set by a org-level bucket configuration, data for new Icechunk Repos are stored within another prefix. By default, the extra prefix is set to Repo name prefixed with 8 random characters. Choose a specific extra prefix by passing the prefix kwarg to create_repo. For example, for a bucket configured with bucket='my-bucket-name' and prefix='my-bucket-prefix',

  1. create_repo("repo-A") stores data in my-bucket-name/my-bucket-prefix/[8-RANDOM_CHARACTERS]_repo_A
  2. create_repo("repo-B", prefix='zoo') stores data in my-bucket-name/my-bucket-prefix/zoo/

Import an existing Icechunk Repo

You can also import an existing Icechunk Repository into Arraylake.

You'll need a bucket configuration for the bucket in which your Icechunk Repository is stored. Let's imagine you created a bucket config with nickname my-bucket, and your bucket contains an Icechunk Repository under the compound prefix my-bucket_prefix/my-icechunk-prefix.

Hit the New Repository button as before, and then open the Additional Settings box. Check the Import Existing Repository button, and type in the prefix under which your Icechunk repo is stored.

Importing an existing Icechunk repository in the web app.

Importing an existing Icechunk repository in the web app.

Repo metadata and descriptions can be set in the same way as when creating a repo from scratch.

If your existing Icechunk Repository contains virtual chunks, this requires an additional configuation step - see the docs on virtual chunks.

Import a Zarr store

You can create a new Icechunk Repository from an existing native Zarr store without copying any data. Arraylake scans the source store and records virtual chunk references that point at the original Zarr chunks in place, then commits them to a brand-new Icechunk Repository. The original store is left untouched, and no chunk data is duplicated.

You'll need:

  • A bucket configuration for the bucket that holds the source Zarr store. The source bucket must use delegated credentials or anonymous (public) access — HMAC buckets can't be used as a virtual chunk source.
  • A destination bucket configuration for the new Repository. This can be the same bucket as the source, or a different one.

From your organization's home page, open the Create Repository menu and choose Import Zarr store. Then fill in:

  • Repository — the name (and optional description) for the new Repository.
  • Source Zarr store — the bucket configuration and prefix where the existing Zarr store lives.
  • Destination — the bucket configuration (and optional prefix) where the new Icechunk Repository will be created.
Importing a Zarr store in the web app.

Importing a Zarr store in the web app.

Submit the form to start the import. Arraylake runs the ingestion in the background and shows its progress; when it finishes, the new Repository is ready to open like any other. Very large stores may be rejected if they exceed the virtual-chunk limit.

For more on how virtual references work and how access to the source data is authorized, see the docs on virtual chunks.

Open a Repo

If you're working in Python, you can open a Repo and start interacting with your data.

repo = client.get_repo("earthmover/ocean")

List Repos

You can list repos associated with an organization.

To browse repositories, navigate to your org's repositories page.

Viewing a list of your org's repositories in the web app.

Viewing a list of your org's repositories in the web app.

You can also filter repos on repo metadata.

In the web app, you can use the search bar in your org's Repositories page.

Searching the list of your org's repositories in the web app.

Searching the list of your org's repositories in the web app.

Delete a Repo

Finally, we can delete a repo.

warning

Deleting a repo cannot be undone! Use this operation carefully.

Navigate to the Settings tab of your Repo:

The settings tab of a repo.

The settings tab of a repo in the web app.

Scroll to the bottom, and hit Delete Repository.

Deleting a Repository via the web app.

Deleting a Repostory via the web app.