Skip to content

Move private-cluster integration test to 1ES pipeline - #549

Open
Suneha Bose (bosesuneha) wants to merge 3 commits into
Azure:mainfrom
bosesuneha:move-pvt-intg-to-1es
Open

Suneha Bose (bosesuneha) wants to merge 3 commits into
Azure:mainfrom
bosesuneha:move-pvt-intg-to-1es

Conversation

@bosesuneha

@bosesuneha Suneha Bose (bosesuneha) commented Jul 20, 2026 •

Copy link
Copy Markdown
Member

GitHub-repo Federated Identity Credentials are being removed, so the private-cluster integration test moves from GitHub Actions to a release-gated Azure DevOps 1ES pipeline authenticated via a governed WIF service connection.

What changes

  • Add .azure-pipelines/1es-integration-tests-private.yml (extends the 1ES Unofficial template; runs on staging-pool-amd64-mariner-2).
  • Delete the GitHub workflow whose repo-FIC-backed azure/login no longer authenticates.
  • Invoke the Node 24 action via node lib/index.js with explicit INPUT_* values and the GitHub context variables the action consumes.
  • Use a pre-created, shared resource group; create/delete only a tagged, per-build cluster (no az group create/delete).
  • Add optional resourceGroup= to k8s-deploy-test.py (defaults to the cluster name) so verification targets the shared resource group.

Release-gate safety

The new 1ES definition deliberately declares:

trigger: none
pr: none

It must not be attached to a PR branch policy or given a CI trigger. The pipeline uses a WIF service connection and therefore must not execute PR-controlled code or provision a private AKS cluster after every merge to main.

Before any Azure-authenticated task runs, the pipeline rejects the run unless all of these are true:

  • the checked-out ref matches refs/heads/releases/vX.Y.Z;
  • the internal release orchestrator supplied a full 40-character releaseCandidateCommit parameter;
  • that parameter equals both Build.SourceVersion and the checked-out HEAD;
  • the release branch version matches package.json.

The internal release orchestrator must queue this pipeline for the exact release-candidate SHA and make tagging/publishing explicitly depend on its success. Merely triggering this pipeline when a release branch is created is not a release gate because publication could proceed concurrently.

Azure DevOps pipeline UI triggers must remain disabled because UI trigger settings can override YAML. The k8s-deploy-intg-test-svc-conn service connection should use a custom role scoped to k8s-deploy-intg-rg, not subscription-wide Contributor or Owner.

Additional hardening

  • Use npm ci rather than npm install.
  • Install a pinned kubectl version explicitly instead of depending on the pool image.
  • Set GITHUB_WORKSPACE, RUNNER_TEMP, commit/ref/run metadata, and action input defaults for parity with GitHub Action execution.
  • Disable image pulling for annotations so the test does not depend on Docker being installed.
  • Use set -euo pipefail for scripts.
  • Tag the temporary cluster with its purpose and build ID.
  • Preserve cleanup with condition: always(), but report deletion failures instead of swallowing them with || true.

Validation

  • YAML parses successfully with js-yaml.
  • npm run format-check passes.
  • npm run typecheck passes.
  • npm test passes: 29 files, 296 tests.
  • npm run build passes.

@bosesuneha
Suneha Bose (bosesuneha) requested a review from a team as a code owner July 20, 2026 22:41
GitHub-repo Federated Identity Credentials are being removed, so the
private-cluster integration test moves from GitHub Actions to an Azure DevOps
1ES pipeline that authenticates via a WIF service connection.

- Add .pipelines/1es-integration-tests-private.yml (extends 1ES Unofficial
  template; runs on staging-pool-amd64-mariner-2).
- Invoke the node24 action via `node lib/index.js` with INPUT_* env vars.
- Use a pre-created, shared resource group; create/delete only a per-build
  cluster (no az group create/delete).
- Add optional resourceGroup= arg to k8s-deploy-test.py (defaults to cluster
  name) so verification targets the shared RG.
- Delete the GitHub workflow whose FIC-based azure/login no longer authenticates.
@bosesuneha

Copy link
Copy Markdown
Member Author

Copilot resolve the merge conflicts in this pull request

@Tatsinnit

Tatsat (Tats) Mishra 🐉 (Tatsinnit) commented Sep 19, 2026 •

Copy link
Copy Markdown
Member

💡 Hiya Suneha Bose (@bosesuneha), I have made this commit, so rather than jsut commenting I will add the commit and collab with you!! The reason, adding context for the latest changes: Commit: ad1fb95

The original pipeline triggers introduced two aspects of interest:

  • A main branch trigger would create a private AKS cluster after every merge, including documentation and dependency-only changes. This adds that anyone whose PR merged can trigger something in internal pipeline, also careful nit security cost to consider and resource usage.
  • A PR trigger targeting releases/* could allow PR-controlled code to execute with the WIF-backed Azure service connection. A pipeline with privileged Azure access should not run from an untrusted or mutable PR context.

A release-branch trigger alone would also not provide a true release gate. The release workflow could continue tagging and publishing while the integration test runs independently.

The pipeline now uses:

trigger: none
pr: none

It is intended to be invoked only by the internal release orchestrator. Before any Azure-authenticated task runs, it verifies that:

  • the source branch matches releases/vX.Y.Z;
  • the orchestrator supplied a full release-candidate commit SHA;
  • the supplied SHA matches both Build.SourceVersion and the checked-out HEAD;
  • the branch version matches the version in package.json.

The internal release pipeline should queue this test for the exact release-candidate commit and make tagging/publishing depend on its successful completion. This ensures the tested code is exactly the code being released.

The update also makes execution more deterministic and secure by using npm ci, explicitly installing a pinned kubectl version, restoring the GitHub Action environment expected by the action, tagging temporary clusters, and surfacing cleanup failures.

Azure DevOps UI triggers should remain disabled because UI trigger settings can override the YAML configuration. The WIF service connection should also use the minimum required role scoped to the shared integration-test resource group rather than subscription-wide Contributor or Owner permissions.

@bosesuneha

Copy link
Copy Markdown
Member Author

Right now the pipeline has trigger: none/pr: none, so nothing actually runs it, and our GitHub release flow doesn't wait on any test before publishing. I'd like to turn this into a real release gate: split the release into a prepare step that just creates the releases/vX.Y.Z branch and a separate publish step, have this pipeline run automatically on releases/*, and post a status back to GitHub from ADO, that publishing waits on. Tatsat (Tats) Mishra 🐉 (@Tatsinnit) would love your thoughts.

@Tatsinnit

Tatsat (Tats) Mishra 🐉 (Tatsinnit) commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Right now the pipeline has trigger: none/pr: none, so nothing actually runs it, and our GitHub release flow doesn't wait on any test before publishing. I'd like to turn this into a real release gate: split the release into a prepare step that just creates the releases/vX.Y.Z branch and a separate publish step, have this pipeline run automatically on releases/*, and post a status back to GitHub from ADO, that publishing waits on.

You're right that, as currently wired, this is not yet an operational release gate: a trusted component still needs to queue the pipeline, and publishing must wait for its result. I agree with splitting the release into prepare and publish phases and reporting the 1ES result back to GitHub.

My concern is specifically with a direct releases/* CI trigger. The intent behind trigger: none / pr: none is to prevent GitHub-controlled branch activity from directly executing code in internal 1ES with the WIF-backed service connection. Because the pipeline definition comes from the triggered revision, checks inside that YAML are not themselves a security boundary if the branch can be modified.

I suggest this flow:

  • Prepare creates releases/vX.Y.Z from an approved commit and records the immutable candidate SHA.
  • A trusted/governed release bridge queues the 1ES pipeline for exactly that SHA and supplies releaseCandidateCommit.
  • 1ES posts its result back to GitHub against that same SHA.
  • Publish waits for that successful result and verifies that the release branch/tag still points to the tested SHA.

This gives us the real release gate you described without making privileged internal 1ES execution generally triggerable by GitHub branch or PR events. If the proposed automatic trigger is a governed bridge using a trusted pipeline definition and an immutable, validated SHA, then we are aligned; I mainly want that trust boundary to remain explicit.

Just some thoughts to share and ideas for security first kind of approach ❤️ what do you think?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants