Skip to content

Migrate from Azure Artifacts

View as Markdown

This page moves the npm, Maven and Python packages you published to Azure Artifacts to CloudRepo and repoints your builds, with Azure Artifacts left running until your builds pass against CloudRepo. It copies package files and nothing else: your feed permissions, upstream sources and views, and every other piece of data that is not a package file, stay in Azure.

  1. List what you have and decide where each feed goes.
  2. Create the repositories in CloudRepo.
  3. Export the files from Azure Artifacts.
  4. Import them into CloudRepo.
  5. Repoint your builds.
  6. Check, then cut over.

You need a CloudRepo organization. If you have none, sign up.

Create a personal access token (PAT) in your Azure DevOps organization, following Microsoft’s Use personal access tokens. Microsoft’s REST pages for listing feeds and packages ask for the vso.packaging scope, “Grants the ability to read feeds and packages”, and its setup pages for npm and Maven create the token with Packaging, Read & write. Choose the narrowest scope that lets the loops below succeed, give the token a short lifetime, and revoke it when the migration is done.

Give the token to curl through a netrc file, so it is on no command line. Create the file empty and readable by you alone first, so the token is never in a file that others can read:

Terminal
touch ~/.azure-artifacts.netrc && chmod 600 ~/.azure-artifacts.netrc

Then put the token in it. Microsoft says an Azure DevOps PAT is accepted as an HTTP Basic password with an empty username, its Maven setup uses your organization’s name as the username, and its Python setup accepts any username with the PAT as the password. The second entry holds the organization’s name:

~/.azure-artifacts.netrc
machine feeds.dev.azure.com
login ""
password AZURE_DEVOPS_PAT
machine pkgs.dev.azure.com
login my-org
password AZURE_DEVOPS_PAT

Then list your feeds with Microsoft’s Get Feeds:

Terminal
SCOPE=my-org
FEEDS_URL=https://feeds.dev.azure.com/$SCOPE
curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get --data-urlencode 'api-version=7.1' \
"$FEEDS_URL/_apis/packaging/feeds" | jq -r '.value[] | [.name, (.project.name // "")] | @tsv'

Expected: one line for each feed, with the name of its project when it has one. Microsoft says the project must be in the URL when a feed was created in a project. For such a feed, set SCOPE=my-org/my-project and use it in every block that follows. Microsoft’s responses carry their results in a value list, as its REST introduction shows.

Microsoft’s What are upstream sources? says that Azure Artifacts saves a copy of a package installed from an upstream source, and that saved packages “remain in the feed even if the upstream source is disabled or removed”. Get Feeds returns the upstreamSources of each feed, and its includeDeletedUpstreams parameter adds the ones that were removed. List those of one feed into upstreams.tsv:

Terminal
SCOPE=my-org
FEEDS_URL=https://feeds.dev.azure.com/$SCOPE
FEED=my-feed
rm -f upstreams.tsv
feeds=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get \
--data-urlencode 'api-version=7.1' --data-urlencode 'includeDeletedUpstreams=true' \
"$FEEDS_URL/_apis/packaging/feeds") \
&& printf '%s' "$feeds" | jq -r --arg feed "$FEED" '.value[] | select(.name == $feed) | .upstreamSources[]?
| [.id, .name, .upstreamSourceType, .location, (.serviceEndpointId // ""), (.deletedDate // "")] | @tsv' > upstreams.tsv \
&& cat upstreams.tsv \
|| { rm -f upstreams.tsv; echo 'upstreams.tsv was not written: fix the error above and run this block again.' >&2; }

Expected: one line for each upstream source of the feed, with its id, name, type, location, the id of the service endpoint that holds the credentials Azure uses to reach it (empty when there is none), and the date it was removed (empty when it was not). Microsoft’s Get Feeds page gives the type public as a “Publicly available source” and internal as an “Azure DevOps upstream source”. No line means the feed has no upstream sources. An error message means upstreams.tsv was not written: fix it and run the block again.

Then list the packages in one feed with Get Packages, every version of each, into versions.tsv. Microsoft pages the result with $top and $skip, so the function reads pages until one is empty, and stops at the first error:

Terminal
SCOPE=my-org
FEEDS_URL=https://feeds.dev.azure.com/$SCOPE
FEED=my-feed
list_versions() {
local skip=0 page n
while :; do
page=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get \
--data-urlencode 'api-version=7.1' --data-urlencode 'includeAllVersions=true' \
--data-urlencode '$top=100' --data-urlencode "\$skip=$skip" \
"$FEEDS_URL/_apis/packaging/Feeds/$FEED/packages") || return 1
n=$(printf '%s' "$page" | jq '.value | length') || return 1
[ "$n" -gt 0 ] || return 0
printf '%s' "$page" | jq -r '.value[] | . as $p | $p.versions[]?
| [$p.protocolType, $p.name, $p.id, .version, .id, (.directUpstreamSourceId // "")] | @tsv' || return 1
skip=$((skip + n))
done
}
if list_versions > versions.tsv; then
echo 'Versions by package type:'
awk -F'\t' '{ n[$1]++ } END { for (t in n) print n[t], t }' versions.tsv | sort
echo 'Versions saved from an upstream source, by upstream id:'
awk -F'\t' 'FILENAME == ARGV[1] { known[$1] = $2 " (" $3 ", " $4 ")"; next }
$6 != "" { n[$6]++ }
END { for (u in n) print n[u], u, ((u in known) ? known[u] : "(not in upstreams.tsv)") }' upstreams.tsv versions.tsv | sort
else
rm -f versions.tsv
echo 'The list stopped at the error above, so versions.tsv was removed. Run this block again.' >&2
fi

Expected: versions.tsv has one line for each package version, with its package type, name, package id, version, version id and, for a version saved from an upstream source, the id of that upstream. Microsoft’s Get Packages page describes that field, directUpstreamSourceId, as the “upstream source this package was ingested from”. The first count is the versions of each package type. The second is the versions saved from each upstream, next to what upstreams.tsv says that upstream is. If an error is printed above them, the list is incomplete and versions.tsv is gone: nothing below reads a partial list.

A proxy repository fills itself from its upstream when a build asks, so a version saved from an upstream that CloudRepo proxies needs no export. Name an upstream’s id in PROXIED only when both of these hold: the Supported Remote Servers list has that upstream (compare its location in upstreams.tsv), and the service endpoint column is empty, because a Maven, npm or Python proxy repository presents no credential to its upstream (see Upstream Credentials). Leave out of PROXIED every other upstream: another Azure DevOps feed (type internal), a registry CloudRepo does not offer, an id that upstreams.tsv does not show, and an upstream that no longer serves a version you need. No CloudRepo proxy repository can fetch those versions, so they are exported with the rest.

Terminal
PROXIED='' # upstream ids from upstreams.tsv, separated by spaces, for example PROXIED='id1 id2'
: > packages.tsv; : > left-out.tsv
awk -F'\t' -v px="$PROXIED" '
BEGIN { n = split(px, a, " "); for (i = 1; i <= n; i++) proxied[a[i]] = 1 }
$6 in proxied { print > "left-out.tsv"; next }
{ print > "packages.tsv" }' versions.tsv
echo 'To export, by package type:'
awk -F'\t' '{ n[$1]++ } END { for (t in n) print n[t], t }' packages.tsv | sort
echo 'Left to a proxy repository, by upstream id:'
awk -F'\t' '{ n[$6]++ } END { for (u in n) print n[u], u }' left-out.tsv | sort

Expected: packages.tsv has the versions to export, and left-out.tsv the versions saved from the upstreams you named. With PROXIED empty nothing is left out, and every version is exported.

Then decide where each package type goes. The left column uses the package types Microsoft lists in What is Azure Artifacts?:

In Azure Artifacts In CloudRepo
npm, Maven or Python packages published to a feed A repository of that format. Its files are what you export and import. Gradle uses a Maven repository.
A version saved from an upstream source, such as npmjs.com or Maven Central A proxy repository when CloudRepo offers that upstream and it needs no credential. Nothing to export for these versions: a proxy fills itself from the upstream. A version saved from any other upstream is exported like a published one.
Any other package type (NuGet, Cargo, Universal Packages) No CloudRepo repository for it. CloudRepo hosts Maven, Python, npm and Docker.

For each Azure Artifacts feed you are moving, create a repository of the same format in CloudRepo for each package type the feed holds, and give it the name your builds already use if you can. See Creating a repository. Then add whoever needs access: add users, and create a repository token with Read + write for the import and Read only tokens for builds that only pull.

Export one package type at a time from the packages.tsv of step 1. Each loop writes under export/, and the import loops read it from there.

Give npm the token through its own configuration file, as Microsoft’s npm setup shows: a username (any value that is not empty), the PAT Base64-encoded as _password, and an email that npm requires and does not use, for the feed’s registry URL. This block reads the PAT without putting it on a command line and writes the file for owner access only:

Terminal
SCOPE=my-org
FEED=my-feed
read -r -s -p 'Azure DevOps PAT: ' AZURE_DEVOPS_PAT; echo
B64=$(printf '%s' "$AZURE_DEVOPS_PAT" | base64 | tr -d '\n')
( umask 077
for prefix in registry/ ''; do
printf '//pkgs.dev.azure.com/%s/_packaging/%s/npm/%s:username=%s\n' "$SCOPE" "$FEED" "$prefix" my-org
printf '//pkgs.dev.azure.com/%s/_packaging/%s/npm/%s:_password=%s\n' "$SCOPE" "$FEED" "$prefix" "$B64"
printf '//pkgs.dev.azure.com/%s/_packaging/%s/npm/%s:email=npm requires an email and does not use it\n' "$SCOPE" "$FEED" "$prefix"
done > azure.npmrc )
unset AZURE_DEVOPS_PAT B64

Then fetch each version with npm pack, which, for a name and a version, “will fetch it to the cache, copy the tarball to the current working directory as <name>-<version>.tgz” (here, --pack-destination names the directory), pointing it at the feed’s registry URL with the file above as its user configuration:

Terminal
SCOPE=my-org
FEED=my-feed
REGISTRY=https://pkgs.dev.azure.com/$SCOPE/_packaging/$FEED/npm/registry/
mkdir -p export/npm
awk -F'\t' 'tolower($1) == "npm" { print $2 "\t" $4 }' packages.tsv | while IFS=$'\t' read -r name ver; do
npm pack "$name@$ver" --registry "$REGISTRY" --userconfig "$PWD/azure.npmrc" --pack-destination export/npm > /dev/null
done

Expected: export/npm/ holds one .tgz for each npm version in packages.tsv. Delete azure.npmrc when you are done: it holds your PAT.

A Maven feed’s address is a Maven repository: Microsoft’s Maven setup lists https://pkgs.dev.azure.com/<ORGANIZATION_NAME>/<PROJECT_NAME>/_packaging/<FEED_NAME>/maven/v1 under a project’s <repositories>, and Maven’s repository layout puts a file at the groupId as a directory, the artifactId, the version and a name of the form <artifactId>-<version>.<extension>, or <artifactId>-<version>-<classifier>.<extension>, under the repository’s address. Microsoft’s Get Package Version returns the files of a version, for the package types that hold several files in one, which is where the file names come from.

Microsoft’s Get Packages page calls a package’s name only “the display name of the package”, and this loop splits a Maven package’s name at the colon into its groupId and artifactId. Print a few names (awk -F'\t' 'tolower($1) == "maven" { print $2 }' packages.tsv | sort -u | head) and confirm that they read groupId:artifactId for your feed. If they do not, build the group and artifact from your poms instead.

Terminal
SCOPE=my-org
FEEDS_URL=https://feeds.dev.azure.com/$SCOPE
FEED=my-feed
MAVEN_URL=https://pkgs.dev.azure.com/$SCOPE/_packaging/$FEED/maven/v1
awk -F'\t' 'tolower($1) == "maven" { print $2 "\t" $3 "\t" $4 "\t" $5 }' packages.tsv \
| while IFS=$'\t' read -r name pid ver vid; do
group=${name%%:*}
artifact=${name#*:}
path="$(printf '%s' "$group" | tr . /)/$artifact/$ver"
detail=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get \
--data-urlencode 'api-version=7.1' \
"$FEEDS_URL/_apis/packaging/Feeds/$FEED/Packages/$pid/versions/$vid") || continue
files=$(printf '%s' "$detail" | jq -r '.files[]?.name')
[ -n "$files" ] || echo "no files listed for $name $ver" >&2
printf '%s\n' "$files" | while IFS= read -r file; do
[ -n "$file" ] || continue
curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --create-dirs \
--output "export/maven/$path/$file" "$MAVEN_URL/$path/$file"
done
done

Expected: export/maven/<groupId as a path>/<artifactId>/<version>/ holds each version’s files, the layout CloudRepo takes as it is, and the loop prints nothing. A line no files listed means Azure returned no file names for that version: fetch its files from the package’s page in Azure DevOps. An error from the first curl, which prints it, means that version was skipped and none of its files are in export/: run the loop again. An error from the second curl, which prints it, means a file name that Azure listed is not at that path: stop and check the name and the path for that package.

A feed’s Python packages are served as a PyPI index. Microsoft’s Python page gives its address, https://pkgs.dev.azure.com/<ORGANIZATION_NAME>/<PROJECT_NAME>/_packaging/<FEED_NAME>/pypi/simple/, and its Consume packages from PyPI page says to “Enter any value for the username, and use your PAT as the password”, and shows the index address with a user name and the PAT before the host name. The pkgs.dev.azure.com entry of your netrc file, which holds your organization’s name and the PAT, is a credential that index takes. A PEP 503 index lists each package’s files on one page, as links whose text is the file name. The loop below asks Get Package Version for the files of each version, as the Maven loop does, finds each file’s link on its package’s index page and fetches it. It fetches every file Azure lists, which pip download would not: pip saves only the one file it would install on the machine that runs it, so a version held as several wheels, or as wheels and a source distribution, would give you one of them. It needs python3, to read the index page:

Terminal
SCOPE=my-org
FEEDS_URL=https://feeds.dev.azure.com/$SCOPE
FEED=my-feed
PYPI_URL=https://pkgs.dev.azure.com/$SCOPE/_packaging/$FEED/pypi/simple
awk -F'\t' 'tolower($1) == "pypi" { print $2 "\t" $3 "\t" $4 "\t" $5 }' packages.tsv \
| while IFS=$'\t' read -r name pid ver vid; do
norm=$(printf '%s' "$name" | tr 'A-Z' 'a-z' | sed -E 's/[-_.]+/-/g')
detail=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --get \
--data-urlencode 'api-version=7.1' \
"$FEEDS_URL/_apis/packaging/Feeds/$FEED/Packages/$pid/versions/$vid") \
|| { echo "not exported: $name $ver" >&2; continue; }
files=$(printf '%s' "$detail" | jq -r '.files[]?.name')
[ -n "$files" ] || echo "no files listed for $name $ver" >&2
page=$(curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail \
"$PYPI_URL/$norm/") || { echo "not exported: $name $ver" >&2; continue; }
links=$(printf '%s' "$page" | python3 -I -c '
import sys, urllib.parse
from html.parser import HTMLParser
class Links(HTMLParser):
href = None
def handle_starttag(self, tag, attrs):
self.href = dict(attrs).get("href") if tag == "a" else None
def handle_data(self, data):
if self.href:
print(data.strip(), urllib.parse.urljoin(sys.argv[1], self.href).split("#")[0])
self.href = None
Links().feed(sys.stdin.read())' "$PYPI_URL/$norm/")
printf '%s\n' "$files" | while IFS= read -r file; do
[ -n "$file" ] || continue
url=$(printf '%s\n' "$links" | awk -v f="$file" '$1 == f { print $2; exit }')
[ -n "$url" ] || { echo "no link for $file on the index page of $name" >&2; continue; }
curl --netrc-file ~/.azure-artifacts.netrc --silent --show-error --fail --proto =https --location --create-dirs \
--output "export/python/$file" "$url" || { rm -f "export/python/$file"; echo "not exported: $file" >&2; }
done
done

Expected: export/python/ holds each file that Azure lists for each Python version in packages.tsv, and the loop prints nothing. Each line it prints names what is not in export/python/. no files listed means Azure returned no file names for that version, no link for means the package’s index page does not link a file that Azure listed, and not exported follows the error that curl printed for that version or file. Fetch what a line names from the package’s page in Azure DevOps, or run the loop again once you have fixed the cause: a 401 means Azure refused the token. Microsoft’s Download Package page for a Python file says the API “is intended for manual UI download options, not for programmatic access and scripting”, which is why the loop reads the index instead.

Import artifacts into CloudRepo has one loop for each format. Run it for each repository, then check the bytes arrived. You can run a loop again after a failure, and what a second run skips and what it replaces depends on the format: see Running a loop again.

In each build, replace the Azure Artifacts feed URL and credential with the CloudRepo repository’s URL and a repository token. The page for your build tool shows the file to change:

Pull Maven artifacts and Publish Maven artifacts: ~/.m2/settings.xml holds the credential, and pom.xml holds the repository URL.

The username for every client except npm is the email address of the account that created the token, and the password is the token. In CI, store the token in the CI system’s secret store and expose it as an environment variable, as the pages above show; do not write it into a script or a committed file.

Do this in order, and leave Azure Artifacts running until the last step:

  1. Build against CloudRepo, with Azure Artifacts still up. Point one project at CloudRepo and run its full build, then the projects that depend on it. A 401, 403 or 404 has the same causes as when you publish by hand: see “When a Client Answers 401, 403 or 404” on Repository tokens.
  2. Publish one new version to CloudRepo from CI and install it from a second machine.
  3. Switch every build that still names Azure Artifacts, and every CI job’s secret, to CloudRepo.
  4. Stop publishing to Azure Artifacts. From now on a new version exists in CloudRepo only.
  5. Keep Azure Artifacts, or an export of it, until you are sure nothing reads it. The export of step 3 holds the npm, Maven and Python versions in packages.tsv: those published to the feed, and those saved from an upstream you did not name in PROXIED. Whatever a Maven or Python loop printed a line about is missing from export/: fetch it from Azure DevOps before you delete the feed. The export does not hold the versions in left-out.tsv, which a CloudRepo proxy repository fetches from its upstream, nor a feed’s permissions or views, nor NuGet, Cargo or Universal packages. Check those before you delete the feed. Revoke the Azure DevOps token you made for the export, and delete azure.npmrc and the netrc file.

If you are stuck, email support@cloudrepo.io with your organization name, the repositories involved, the command you ran and the error it printed.