Migrate from Azure Artifacts
This page moves the npm, Maven and Python packages you published to Azure Artifacts to CloudRepo and repoints your builds, with Azure Artifacts left running until your builds pass against CloudRepo. It copies package files and nothing else: your feed permissions, upstream sources and views, and every other piece of data that is not a package file, stay in Azure.
- List what you have and decide where each feed goes.
- Create the repositories in CloudRepo.
- Export the files from Azure Artifacts.
- Import them into CloudRepo.
- Repoint your builds.
- Check, then cut over.
1. List what you have
Section titled “1. List what you have”You need a CloudRepo organization. If you have none, sign up.
Create a personal access token (PAT) in your Azure DevOps organization, following Microsoft’s
Use personal access tokens.
Microsoft’s REST pages for listing feeds and packages ask for the vso.packaging scope, “Grants the ability to read
feeds and packages”, and its setup pages for npm
and Maven create the token with
Packaging, Read & write. Choose the narrowest scope that lets the loops below succeed, give the token a short
lifetime, and revoke it when the migration is done.
Give the token to curl through a netrc file, so it is on no command line. Create the file empty and readable by you alone first, so the token is never in a file that others can read:
touch ~/.azure-artifacts.netrc && chmod 600 ~/.azure-artifacts.netrcThen put the token in it. Microsoft says an Azure DevOps PAT is accepted as an HTTP Basic password with an empty username, its Maven setup uses your organization’s name as the username, and its Python setup accepts any username with the PAT as the password. The second entry holds the organization’s name:
machine feeds.dev.azure.comlogin ""password AZURE_DEVOPS_PAT
machine pkgs.dev.azure.comlogin my-orgpassword AZURE_DEVOPS_PATThen list your feeds with Microsoft’s Get Feeds:
SCOPE=my-orgFEEDS_URL=https://feeds.dev.azure.com/$SCOPEcurl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get --data-urlencode 'api-version=7.1' \ "$FEEDS_URL/_apis/packaging/feeds" | jq -r '.value[] | [.name, (.project.name // "")] | @tsv'Expected: one line for each feed, with the name of its project when it has one. Microsoft says the project must be in the
URL when a feed was created in a project. For such a feed, set SCOPE=my-org/my-project and use it in every block that follows.
Microsoft’s responses carry their results in a value list, as its
REST introduction shows.
Microsoft’s What are upstream sources?
says that Azure Artifacts saves a copy of a package installed from an upstream source, and that saved packages “remain in the feed even
if the upstream source is disabled or removed”. Get Feeds returns the upstreamSources of each feed, and its includeDeletedUpstreams
parameter adds the ones that were removed. List those of one feed into upstreams.tsv:
SCOPE=my-orgFEEDS_URL=https://feeds.dev.azure.com/$SCOPEFEED=my-feedrm -f upstreams.tsvfeeds=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get \ --data-urlencode 'api-version=7.1' --data-urlencode 'includeDeletedUpstreams=true' \ "$FEEDS_URL/_apis/packaging/feeds") \ && printf '%s' "$feeds" | jq -r --arg feed "$FEED" '.value[] | select(.name == $feed) | .upstreamSources[]? | [.id, .name, .upstreamSourceType, .location, (.serviceEndpointId // ""), (.deletedDate // "")] | @tsv' > upstreams.tsv \ && cat upstreams.tsv \ || { rm -f upstreams.tsv; echo 'upstreams.tsv was not written: fix the error above and run this block again.' >&2; }Expected: one line for each upstream source of the feed, with its id, name, type, location, the id of the service endpoint that
holds the credentials Azure uses to reach it (empty when there is none), and the date it was removed (empty when it was not).
Microsoft’s Get Feeds page gives the type public as a “Publicly available source” and internal as an “Azure DevOps upstream source”. No line
means the feed has no upstream sources. An error message means upstreams.tsv was not written: fix it and run the block again.
Then list the packages in one feed with Get Packages,
every version of each, into versions.tsv. Microsoft pages the result with $top and $skip, so the function reads pages
until one is empty, and stops at the first error:
SCOPE=my-orgFEEDS_URL=https://feeds.dev.azure.com/$SCOPEFEED=my-feedlist_versions() { local skip=0 page n while :; do page=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get \ --data-urlencode 'api-version=7.1' --data-urlencode 'includeAllVersions=true' \ --data-urlencode '$top=100' --data-urlencode "\$skip=$skip" \ "$FEEDS_URL/_apis/packaging/Feeds/$FEED/packages") || return 1 n=$(printf '%s' "$page" | jq '.value | length') || return 1 [ "$n" -gt 0 ] || return 0 printf '%s' "$page" | jq -r '.value[] | . as $p | $p.versions[]? | [$p.protocolType, $p.name, $p.id, .version, .id, (.directUpstreamSourceId // "")] | @tsv' || return 1 skip=$((skip + n)) done}if list_versions > versions.tsv; then echo 'Versions by package type:' awk -F'\t' '{ n[$1]++ } END { for (t in n) print n[t], t }' versions.tsv | sort echo 'Versions saved from an upstream source, by upstream id:' awk -F'\t' 'FILENAME == ARGV[1] { known[$1] = $2 " (" $3 ", " $4 ")"; next } $6 != "" { n[$6]++ } END { for (u in n) print n[u], u, ((u in known) ? known[u] : "(not in upstreams.tsv)") }' upstreams.tsv versions.tsv | sortelse rm -f versions.tsv echo 'The list stopped at the error above, so versions.tsv was removed. Run this block again.' >&2fiExpected: versions.tsv has one line for each package version, with its package type, name, package id, version, version id and, for a
version saved from an upstream source, the id of that upstream. Microsoft’s Get Packages page describes that field,
directUpstreamSourceId, as the “upstream source this package was ingested from”. The first count is the versions of each package
type. The second is the versions saved from each upstream, next to what upstreams.tsv says that upstream is. If an error
is printed above them, the list is incomplete and versions.tsv is gone: nothing below reads a partial list.
A proxy repository fills itself from its upstream when a build asks, so a version saved from
an upstream that CloudRepo proxies needs no export. Name an upstream’s id in PROXIED only when both of these hold: the
Supported Remote Servers list has that upstream (compare its
location in upstreams.tsv), and the service endpoint column is empty, because a Maven, npm or Python proxy repository presents no credential to its
upstream (see Upstream Credentials). Leave out of PROXIED every other upstream: another
Azure DevOps feed (type internal), a registry CloudRepo does not offer, an id that upstreams.tsv does not show, and an upstream that no longer serves a version you
need. No CloudRepo proxy repository can fetch those versions, so they are exported with the rest.
PROXIED='' # upstream ids from upstreams.tsv, separated by spaces, for example PROXIED='id1 id2': > packages.tsv; : > left-out.tsvawk -F'\t' -v px="$PROXIED" ' BEGIN { n = split(px, a, " "); for (i = 1; i <= n; i++) proxied[a[i]] = 1 } $6 in proxied { print > "left-out.tsv"; next } { print > "packages.tsv" }' versions.tsvecho 'To export, by package type:'awk -F'\t' '{ n[$1]++ } END { for (t in n) print n[t], t }' packages.tsv | sortecho 'Left to a proxy repository, by upstream id:'awk -F'\t' '{ n[$6]++ } END { for (u in n) print n[u], u }' left-out.tsv | sortExpected: packages.tsv has the versions to export, and left-out.tsv the versions saved from the upstreams you named. With
PROXIED empty nothing is left out, and every version is exported.
Then decide where each package type goes. The left column uses the package types Microsoft lists in What is Azure Artifacts?:
| In Azure Artifacts | In CloudRepo |
|---|---|
| npm, Maven or Python packages published to a feed | A repository of that format. Its files are what you export and import. Gradle uses a Maven repository. |
| A version saved from an upstream source, such as npmjs.com or Maven Central | A proxy repository when CloudRepo offers that upstream and it needs no credential. Nothing to export for these versions: a proxy fills itself from the upstream. A version saved from any other upstream is exported like a published one. |
| Any other package type (NuGet, Cargo, Universal Packages) | No CloudRepo repository for it. CloudRepo hosts Maven, Python, npm and Docker. |
2. Create the repositories
Section titled “2. Create the repositories”For each Azure Artifacts feed you are moving, create a repository of the same format in CloudRepo for each package type the feed holds, and give it the name your builds already use if you can. See Creating a repository. Then add whoever needs access: add users, and create a repository token with Read + write for the import and Read only tokens for builds that only pull.
3. Export the files
Section titled “3. Export the files”Export one package type at a time from the packages.tsv of step 1. Each loop writes under export/, and the
import loops read it from there.
Give npm the token through its own configuration file, as Microsoft’s
npm setup shows: a username (any value that is not empty), the
PAT Base64-encoded as _password, and an email that npm requires and does not use, for the feed’s registry URL. This block reads the PAT
without putting it on a command line and writes the file for owner access only:
SCOPE=my-orgFEED=my-feedread -r -s -p 'Azure DevOps PAT: ' AZURE_DEVOPS_PAT; echoB64=$(printf '%s' "$AZURE_DEVOPS_PAT" | base64 | tr -d '\n')( umask 077 for prefix in registry/ ''; do printf '//pkgs.dev.azure.com/%s/_packaging/%s/npm/%s:username=%s\n' "$SCOPE" "$FEED" "$prefix" my-org printf '//pkgs.dev.azure.com/%s/_packaging/%s/npm/%s:_password=%s\n' "$SCOPE" "$FEED" "$prefix" "$B64" printf '//pkgs.dev.azure.com/%s/_packaging/%s/npm/%s:email=npm requires an email and does not use it\n' "$SCOPE" "$FEED" "$prefix" done > azure.npmrc )unset AZURE_DEVOPS_PAT B64Then fetch each version with npm pack, which, for a name and a version, “will fetch it to the cache, copy the tarball to the
current working directory as <name>-<version>.tgz” (here, --pack-destination names the directory), pointing it at the feed’s registry URL with the file above as its user configuration:
SCOPE=my-orgFEED=my-feedREGISTRY=https://pkgs.dev.azure.com/$SCOPE/_packaging/$FEED/npm/registry/mkdir -p export/npmawk -F'\t' 'tolower($1) == "npm" { print $2 "\t" $4 }' packages.tsv | while IFS=$'\t' read -r name ver; do npm pack "$name@$ver" --registry "$REGISTRY" --userconfig "$PWD/azure.npmrc" --pack-destination export/npm > /dev/nulldoneExpected: export/npm/ holds one .tgz for each npm version in packages.tsv. Delete azure.npmrc when you are done: it holds your PAT.
A Maven feed’s address is a Maven repository: Microsoft’s Maven setup lists
https://pkgs.dev.azure.com/<ORGANIZATION_NAME>/<PROJECT_NAME>/_packaging/<FEED_NAME>/maven/v1 under a project’s <repositories>, and Maven’s
repository layout puts a file at the groupId as a directory, the artifactId, the version and a name of the form
<artifactId>-<version>.<extension>, or <artifactId>-<version>-<classifier>.<extension>, under the repository’s address. Microsoft’s
Get Package Version returns the
files of a version, for the package types that hold several files in one, which is where the file names come from.
Microsoft’s Get Packages page calls a package’s name only “the display name of the package”, and this loop splits a Maven package’s
name at the colon into its groupId and artifactId. Print a few names (awk -F'\t' 'tolower($1) == "maven" { print $2 }' packages.tsv | sort -u | head) and confirm that
they read groupId:artifactId for your feed. If they do not, build the group and artifact from your poms instead.
SCOPE=my-orgFEEDS_URL=https://feeds.dev.azure.com/$SCOPEFEED=my-feedMAVEN_URL=https://pkgs.dev.azure.com/$SCOPE/_packaging/$FEED/maven/v1awk -F'\t' 'tolower($1) == "maven" { print $2 "\t" $3 "\t" $4 "\t" $5 }' packages.tsv \ | while IFS=$'\t' read -r name pid ver vid; do group=${name%%:*} artifact=${name#*:} path="$(printf '%s' "$group" | tr . /)/$artifact/$ver" detail=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get \ --data-urlencode 'api-version=7.1' \ "$FEEDS_URL/_apis/packaging/Feeds/$FEED/Packages/$pid/versions/$vid") || continue files=$(printf '%s' "$detail" | jq -r '.files[]?.name') [ -n "$files" ] || echo "no files listed for $name $ver" >&2 printf '%s\n' "$files" | while IFS= read -r file; do [ -n "$file" ] || continue curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --create-dirs \ --output "export/maven/$path/$file" "$MAVEN_URL/$path/$file" done doneExpected: export/maven/<groupId as a path>/<artifactId>/<version>/ holds each version’s files, the layout CloudRepo takes as it is, and
the loop prints nothing. A line no files listed means Azure returned no file names for that version: fetch its files from the
package’s page in Azure DevOps. An error from the first curl, which prints it, means that version was skipped and none of its files
are in export/: run the loop again. An error from the second curl, which prints it, means a file name that Azure listed is not at that path: stop and check the
name and the path for that package.
Python
Section titled “Python”A feed’s Python packages are served as a PyPI index. Microsoft’s Python page
gives its address, https://pkgs.dev.azure.com/<ORGANIZATION_NAME>/<PROJECT_NAME>/_packaging/<FEED_NAME>/pypi/simple/, and its
Consume packages from PyPI page says to
“Enter any value for the username, and use your PAT as the password”, and shows the index address with a user name and the PAT before the host name. The
pkgs.dev.azure.com entry of your netrc file, which holds your organization’s name and the PAT, is a credential that index takes.
A PEP 503 index lists each package’s files on one page, as links whose text is the file name. The loop below asks Get Package Version for the
files of each version, as the Maven loop does, finds each file’s link on its package’s index page and fetches it. It fetches every file Azure lists, which pip download would not: pip saves
only the one file it would install on the machine that runs it, so a version held as several wheels, or as wheels and a source
distribution, would give you one of them. It needs python3, to read the index page:
SCOPE=my-orgFEEDS_URL=https://feeds.dev.azure.com/$SCOPEFEED=my-feedPYPI_URL=https://pkgs.dev.azure.com/$SCOPE/_packaging/$FEED/pypi/simpleawk -F'\t' 'tolower($1) == "pypi" { print $2 "\t" $3 "\t" $4 "\t" $5 }' packages.tsv \ | while IFS=$'\t' read -r name pid ver vid; do norm=$(printf '%s' "$name" | tr 'A-Z' 'a-z' | sed -E 's/[-_.]+/-/g') detail=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get \ --data-urlencode 'api-version=7.1' \ "$FEEDS_URL/_apis/packaging/Feeds/$FEED/Packages/$pid/versions/$vid") \ || { echo "not exported: $name $ver" >&2; continue; } files=$(printf '%s' "$detail" | jq -r '.files[]?.name') [ -n "$files" ] || echo "no files listed for $name $ver" >&2 page=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail \ "$PYPI_URL/$norm/") || { echo "not exported: $name $ver" >&2; continue; } links=$(printf '%s' "$page" | python3 -I -c 'import sys, urllib.parsefrom html.parser import HTMLParserclass Links(HTMLParser): href = None def handle_starttag(self, tag, attrs): self.href = dict(attrs).get("href") if tag == "a" else None def handle_data(self, data): if self.href: print(data.strip(), urllib.parse.urljoin(sys.argv[1], self.href).split("#")[0]) self.href = NoneLinks().feed(sys.stdin.read())' "$PYPI_URL/$norm/") printf '%s\n' "$files" | while IFS= read -r file; do [ -n "$file" ] || continue url=$(printf '%s\n' "$links" | awk -v f="$file" '$1 == f { print $2; exit }') [ -n "$url" ] || { echo "no link for $file on the index page of $name" >&2; continue; } curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --proto =https --location --create-dirs \ --output "export/python/$file" "$url" || { rm -f "export/python/$file"; echo "not exported: $file" >&2; } done doneExpected: export/python/ holds each file that Azure lists for each Python version in packages.tsv, and the loop prints nothing. Each line
it prints names what is not in export/python/. no files listed means Azure returned no file names for that version, no link for means
the package’s index page does not link a file that Azure listed, and not exported follows the error that curl printed for that version or file. Fetch what a line names from the package’s page in Azure DevOps,
or run the loop again once you have fixed the cause: a 401 means Azure refused the token. Microsoft’s
Download Package
page for a Python file says the API “is intended for manual UI download options, not for programmatic access and scripting”, which is why the loop reads the index instead.
4. Import the files
Section titled “4. Import the files”Import artifacts into CloudRepo has one loop for each format. Run it for each repository, then check the bytes arrived. You can run a loop again after a failure, and what a second run skips and what it replaces depends on the format: see Running a loop again.
5. Repoint your builds
Section titled “5. Repoint your builds”In each build, replace the Azure Artifacts feed URL and credential with the CloudRepo repository’s URL and a repository token. The page for your build tool shows the file to change:
Pull Maven artifacts and Publish Maven artifacts:
~/.m2/settings.xml holds the credential, and pom.xml holds the repository URL.
Pull Gradle dependencies and
Publish Gradle artifacts: ~/.gradle/gradle.properties holds the
credential, and your build script holds the repository URL.
Pull npm packages and Publish npm packages: ~/.npmrc
holds the registry and the token.
Pull Python packages and
Publish Python packages: pip.conf or PIP_INDEX_URL, and ~/.pypirc or
TWINE_* for uploads.
The username for every client except npm is the email address of the account that created the token, and the password is the token. In CI, store the token in the CI system’s secret store and expose it as an environment variable, as the pages above show; do not write it into a script or a committed file.
6. Check, then cut over
Section titled “6. Check, then cut over”Do this in order, and leave Azure Artifacts running until the last step:
- Build against CloudRepo, with Azure Artifacts still up. Point one project at CloudRepo and run its full
build, then the projects that depend on it. A
401,403or404has the same causes as when you publish by hand: see “When a Client Answers 401, 403 or 404” on Repository tokens. - Publish one new version to CloudRepo from CI and install it from a second machine.
- Switch every build that still names Azure Artifacts, and every CI job’s secret, to CloudRepo.
- Stop publishing to Azure Artifacts. From now on a new version exists in CloudRepo only.
- Keep Azure Artifacts, or an export of it, until you are sure nothing reads it. The export of step 3 holds the npm, Maven and
Python versions in
packages.tsv: those published to the feed, and those saved from an upstream you did not name inPROXIED. Whatever a Maven or Python loop printed a line about is missing fromexport/: fetch it from Azure DevOps before you delete the feed. The export does not hold the versions inleft-out.tsv, which a CloudRepo proxy repository fetches from its upstream, nor a feed’s permissions or views, nor NuGet, Cargo or Universal packages. Check those before you delete the feed. Revoke the Azure DevOps token you made for the export, and deleteazure.npmrcand the netrc file.
If you are stuck, email support@cloudrepo.io with your organization name, the repositories involved, the command you ran and the error it printed.