Shrink CI Job Reach: Audit and Reduce Pipeline Identity Access
For developers who run builds, deploys, or migrations from CI and want to know exactly what that job identity can touch. This guide shows how to inspect the effective permissions and network path of a CI job, then reduce both so a compromised pipeline can do less damage.
TL;DR — Your CI job usually has two powers: an identity that can call APIs and a network path that can reach internal services. Most overexposure comes from broad cloud IAM roles, long-lived secrets, and runners placed on flat networks. The highest-value fix is to replace static credentials with short-lived job-issued identity, then scope that identity and the runner's egress separately. Reading time: ~7 min
What it is and where it sits
This is about the blast radius of a CI job if its steps are malicious, compromised, or just wrong. "What your CI job's identity can reach" means two different things that people often mix together:
- Control-plane reach: what APIs the job can call because it has credentials or a federated identity.
- Data-plane reach: what hosts, ports, and services the runner can actually connect to over the network.
You need to shrink both. A job with a narrowly scoped cloud role but unrestricted network egress can still hit internal databases if they trust source IPs or shared passwords. A job on a locked-down subnet but holding an admin API token can still delete infrastructure.
In a typical setup, the CI platform issues a job token or OIDC token to the runner. The runner uses that to fetch code, pull secrets, assume a cloud role, talk to artifact storage, and maybe reach private services for tests or deploys. This replaces the old pattern of stuffing long-lived cloud keys into CI secrets.
Developer push
|
v
Git host / CI control plane
|
| issues job token / OIDC token
v
Runner executes steps
| \
| \__ control plane: STS/IAM, secret manager, artifact registry
|
\__ data plane: DB, cache, internal HTTP service, Kubernetes API, SSH target
Where it lives in the flow:
- Identity lives in the CI platform token service and your cloud/provider IAM trust policy.
- Network reach lives where the runner lives: hosted runner, self-hosted VM, Kubernetes pod, or container on a subnet/VPC.
- Authorization is split across cloud IAM, service-specific ACLs, Kubernetes RBAC, database roles, and firewall/security-group/network-policy rules.
The practical replacement to aim for is: job-scoped, short-lived identity + isolated runner network + service-specific least privilege.
How it actually works
Walk one realistic example: a CI job builds an image, pushes it to a registry, then runs a database migration against a private Postgres instance.
Step 1: The job gets an identity
Modern CI systems can mint an OIDC token per job. The job exchanges that token with your cloud STS endpoint for temporary credentials tied to a role.
The trust policy typically checks claims like repository, branch, workflow name, or audience. If those conditions are loose, any repo or branch can assume the role.
A typical failure mode when trust is wrong looks like this:
$ aws sts assume-role-with-web-identity \
--role-arn arn:aws:iam::123456789012:role/ci-deploy \
--role-session-name ci-test \
--web-identity-token file://oidc-token.jwt
An error occurred (AccessDenied) when calling the AssumeRoleWithWebIdentity operation:
Not authorized to perform sts:AssumeRoleWithWebIdentity
That is good if the wrong job tried it. It is bad if your intended job gets this because your trust conditions do not match the actual token claims.
Step 2: Temporary credentials authorize API actions
Once assumed, the role may allow:
ecr:*or registry push permissionss3:GetObjectfor build inputssecretsmanager:GetSecretValuefor a DB password- maybe
eks:DescribeClusterorecs:UpdateServicefor deploys
This is where many teams overgrant. They give one role both deploy and infrastructure powers because it is convenient.
To inspect what the job can do, ask from inside the job:
aws sts get-caller-identity
aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::123456789012:role/ci-deploy \
--action-names secretsmanager:GetSecretValue rds-db:connect s3:PutObject ecs:UpdateService
If simulate-principal-policy is unavailable in your environment, test the exact API calls you care about and record the result.
Step 3: The runner still needs network path
Even with valid credentials, the migration only works if the runner can reach Postgres on 5432. If the runner is on a public hosted network and the database is private, the TCP connection fails.
Typical output:
$ nc -vz db.internal.example 5432
nc: connect to db.internal.example port 5432 (tcp) failed: Connection timed out
If DNS is wrong or split-horizon DNS is missing:
$ getent hosts db.internal.example
$ echo $?
2
If the service is reachable but TLS or auth is wrong:
$ psql "host=db.internal.example port=5432 dbname=app user=ci_migrator sslmode=require"
psql: error: connection to server at "db.internal.example" (10.20.4.15), port 5432 failed: FATAL: password authentication failed for user "ci_migrator"
Step 4: Shrink both dimensions
For this example, the safer design is:
- Build job role can push images, nothing else.
- Migration job role can read one secret or, better, use database IAM auth if supported.
- Migration job runs on a runner subnet that can reach only the database and required cloud endpoints.
- Database user
ci_migratorcan run schema changes only on the target database, not read unrelated data or create superusers.
End state: if the build step is compromised, it cannot reach the DB at all. If the migration step is compromised, it cannot delete infrastructure or read every secret.
When to use it (and when not to)
| Scenario | Recommendation |
|---|---|
| CI only builds/tests public code and pushes artifacts to one registry | Use short-lived job identity and a dedicated push-only role. You likely do not need private network reach. |
| CI deploys to cloud resources | Split build and deploy identities. Scope deploy role to one environment and one service family. |
| CI runs migrations or integration tests against private services | Use a dedicated runner network segment plus a separate migration/test identity. |
| You currently store long-lived cloud keys in CI secrets | Replace them first. This is the highest-risk, highest-payoff fix. |
| Small project, no private resources, no cloud admin actions from CI | You probably do not need complex network isolation yet; still stop using admin API tokens. |
| You use one self-hosted runner for every repo/team | Fix this now. Per-runner or per-runner-group isolation matters more than polishing IAM statements. |
You probably don't need the full treatment if your pipeline never touches production, private networks, or mutable infrastructure. But even then, static credentials in CI are still worth removing.
Trade-offs
Short-lived federated identity instead of static secrets
- Benefit: no long-lived keys to leak, easy revocation via trust policy changes.
- Cost: more setup in IAM trust relationships, claim matching, and debugging token audiences/subjects.
Separate roles per job type
- Benefit: compromised build job cannot deploy or read secrets.
- Cost: more roles and policy maintenance; pipelines need clearer boundaries.
Self-hosted or private runners for internal reach
- Benefit: can keep databases and internal APIs off the public internet.
- Cost: patching, autoscaling, image hygiene, and network troubleshooting become your problem.
Network egress restrictions
- Benefit: blocks whole classes of exfiltration and accidental access.
- Cost: allowlists drift; package installs, image pulls, and test fixtures often break first.
Service-specific least privilege
- Benefit: even valid CI identity cannot do unrelated damage.
- Cost: every service has its own auth model: IAM, RBAC, DB roles, ACLs, firewall rules.
Environment separation
- Benefit: staging compromise does not imply production compromise.
- Cost: duplicated config, more secrets/roles, and more operational overhead.
In practice
Example 1: Trust only one repo and branch for a deploy role
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::123456789012:oidc-provider/token.actions.githubusercontent.com"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com"
},
"StringLike": {
"token.actions.githubusercontent.com:sub": "repo:acme/api-service:ref:refs/heads/main"
}
}
}
]
}
This trust policy allows only jobs from one repository on main to assume the role. Gotcha: branch, tag, and environment protection claims differ by CI provider and workflow type; inspect a real token before writing conditions or you will either lock yourself out or overgrant.
Example 2: Minimal policy for image push, not full admin
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"ecr:BatchCheckLayerAvailability",
"ecr:CompleteLayerUpload",
"ecr:InitiateLayerUpload",
"ecr:PutImage",
"ecr:UploadLayerPart"
],
"Resource": "arn:aws:ecr:us-east-1:123456789012:repository/api-service"
},
{
"Effect": "Allow",
"Action": "ecr:GetAuthorizationToken",
"Resource": "*"
}
]
}
This is enough for a push-only build job to upload to one repository. Gotcha: some APIs are account-scoped and require Resource: "*"; do not respond by widening every other statement to * out of frustration.
Example 3: Restrict runner egress with Kubernetes NetworkPolicy
⚠️ If your runner is in Kubernetes and you apply a default-deny egress policy without first allowing DNS and artifact/STS endpoints, jobs will fail immediately. Expect image pulls, package installs, and token exchange to break until you add explicit egress rules.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: ci-runner-egress
namespace: ci
spec:
podSelector:
matchLabels:
app: runner
policyTypes:
- Egress
egress:
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53
- to:
- ipBlock:
cidr: 10.20.4.15/32
ports:
- protocol: TCP
port: 5432
This allows runner pods to resolve DNS and reach only one Postgres host. Gotcha: NetworkPolicy behavior depends on your CNI plugin; verify enforcement with a test pod, not assumptions.
Example 4: Fast reachability checks inside a job
set -euxo pipefail
echo "== identity =="
aws sts get-caller-identity
echo "== dns =="
getent hosts db.internal.example || true
echo "== tcp =="
nc -vz -w 3 db.internal.example 5432 || true
echo "== https control plane =="
curl -I -sS --max-time 5 https://sts.amazonaws.com || true
This gives you a quick picture of who the job is, whether DNS works, whether the target port is reachable, and whether outbound HTTPS works. Gotcha: curl -I only proves HTTP(S) egress, not that your IAM trust or service auth is correct.
A healthy curl -I usually looks like:
HTTP/1.1 200 OK
x-amzn-requestid: 12345678-90ab-cdef-1234-567890abcdef
content-type: text/xml
content-length: 123
A proxy or redirect surprise often looks like:
HTTP/1.1 301 Moved Permanently
location: http://proxy.local/login
content-length: 0
If you see redirects to a captive proxy or auth portal from a runner, fix the network path before debugging IAM.
Further reading
- AWS IAM User Guide: "Configuring a role for OIDC federation"
- Kubernetes Documentation: "Network Policies"
- PostgreSQL Documentation: "Client Authentication"
- NIST SP 800-207 Zero Trust Architecture
- OWASP CI/CD Security Cheat Sheet
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI