Skip to content

fix(ISV-7735): increase IIB task and pipeline timeouts (temporary) - #1076

Merged
tomasbakk merged 1 commit into
mainfrom
ISV-7735
Oct 9, 2026
Merged

tomasbakk merged 1 commit into
mainfrom
ISV-7735

Conversation

@tomasbakk

Copy link
Copy Markdown
Contributor

This change increases timeouts for the rleease pipeline and add-bundle-to-index task.
Reason for this: current IIB architecture for community requests allows only 2 requests in the queue, which results in major increase in time needed to handle those requests.
This is only a temporary fix. They will release a new version on around 5 Nov. We should check and then return the values back.

Closes: https://redhat.atlassian.net/browse/ISV-7735

Merge Request Checklists

  • Development is done in feature branches
  • Code changes are submitted as pull request into a primary branch [Provide reason for non-primary branch submissions]
  • Code changes are covered with unit and integration tests.
  • Code passes all automated code tests:
    • Linting
    • Code formatter - Black
    • Security scanners
    • Unit tests
    • Integration tests
  • Code is reviewed by at least 1 team member
  • Pull request is tagged with "risk/good-to-go" label for minor changes

…karound)

Signed-off-by: tbak <tbak@redhat.com>
@qodo-redhat-openshift-ecosystem

Copy link
Copy Markdown

PR Summary by Qodo

Temporarily extend IIB index-build and pipeline timeouts

🐞 Bug fix ⚙️ Configuration changes 🕐 10-20 Minutes

Grey Divider

AI Description

• Extend IIB batch polling and index-bundle task limits to tolerate longer community request queues.
• Extend the community release PipelineRun limit so its longer-running task can finish.
• Mark the polling increase for review after the planned IIB release.
Diagram

graph TD
  Trigger["Community release trigger"] --> Release["Release pipeline"] --> Task["Index bundle task"] --> Poller["IIB batch poller"] --> IIB["IIB service"]
  Hosted["Hosted pipeline"] --> Task
Loading
High-Level Assessment

Keep the coordinated timeout increases as a temporary workaround: the IIB polling, task step, pipeline task, and release PipelineRun limits must accommodate the same delay. A configurable timeout scheme would add scope for a change intended to be reverted after the IIB release.

Files changed (5) +13 / -6

Bug fix (1) +8 / -1
iib.pyExtend default IIB batch polling wait +8/-1

Extend default IIB batch polling wait

• Raises the default wait_for_batch_results timeout from 90 to 210 minutes. Adds a note tying the temporary increase and planned reversion check to ISV-7735; callers using the default include index addition and operator removal.

operatorcert/iib.py

Other (4) +5 / -5
community-release-pipeline-trigger.ymlExtend community release PipelineRun limits +2/-2

Extend community release PipelineRun limits

• Raises the PipelineRun limit from 4h15m to 6h15m and its aggregate tasks limit from 4h5m to 6h5m, leaving room for the longer index build.

ansible/roles/operator-pipeline/tasks/community-release-pipeline-trigger.yml

operator-hosted-pipeline.ymlExtend hosted index-bundle task limit +1/-1

Extend hosted index-bundle task limit

• Raises the hosted pipeline's add-bundle-to-index task timeout from 90 minutes to three hours and 30 minutes.

ansible/roles/operator-pipeline/templates/openshift/pipelines/operator-hosted-pipeline.yml

operator-release-pipeline.ymlExtend release index-bundle task limit +1/-1

Extend release index-bundle task limit

• Raises the release pipeline's add-bundle-to-index task timeout from 90 minutes to three hours and 30 minutes.

ansible/roles/operator-pipeline/templates/openshift/pipelines/operator-release-pipeline.yml

add-bundle-to-index.ymlExtend index-bundle step limit +1/-1

Extend index-bundle step limit

• Raises the shared add-bundle-to-index task step timeout from 90 minutes to three hours and 30 minutes.

ansible/roles/operator-pipeline/templates/openshift/tasks/add-bundle-to-index.yml

@qodo-redhat-openshift-ecosystem

qodo-redhat-openshift-ecosystem Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Hosted bundle updates still time out early ✗ Dismissed
Description
The hosted pipeline now gives add-bundle-to-index 3h30m, but both hosted triggers still cap all
tasks at 2h20m and the pipeline at 2h30m. When a non-FBC bundle update needs more than 2h20m, the
run ends before the new task timeout can help.
Code

ansible/roles/operator-pipeline/templates/openshift/pipelines/operator-hosted-pipeline.yml[911]

+      timeout: "3h30m0s"
Relevance

●●● Strong

Clear timeout mismatch defeats the PR’s hosted bundle-timeout intent and is a deterministic
integration bug.

PR-#996

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The changed hosted task permits 3h30m and applies to non-FBC bundle additions, while each hosted
trigger creates runs with a 2h20m task limit and a 2h30m pipeline limit.

ansible/roles/operator-pipeline/templates/openshift/pipelines/operator-hosted-pipeline.yml[908-918]
ansible/roles/operator-pipeline/tasks/operator-hosted-pipeline-trigger.yml[99-104]
ansible/roles/operator-pipeline/tasks/community-hosted-pipeline-trigger.yml[102-107]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Hosted runs cannot use the new 3h30m bundle-task timeout because their run-level task and pipeline limits are shorter.

## Fix Focus Areas
- ansible/roles/operator-pipeline/templates/openshift/pipelines/operator-hosted-pipeline.yml[908-918]
- ansible/roles/operator-pipeline/tasks/operator-hosted-pipeline-trigger.yml[99-104]
- ansible/roles/operator-pipeline/tasks/community-hosted-pipeline-trigger.yml[102-107]

## Recommended Fix
Increase the task and pipeline limits in both hosted triggers so the bundle task can run for 3h30m, with enough additional pipeline time for the remaining tasks.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Index removals are killed while polling ✗ Dismissed
Description
rm_operator_from_index inherits the new 210-minute default from wait_for_batch_results, but its
execution step still has a 90-minute timeout. If an index-removal build remains pending at that
boundary, the step terminates while the poller is still waiting rather than reaching its own
timeout.
Code

operatorcert/iib.py[164]

+    iib_url: str, batch_id: int, timeout: float = 210 * 60, delay: float = 20
Relevance

●●● Strong

The shorter execution timeout deterministically kills polling before the newly extended poller
timeout can take effect.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The removal caller supplies no timeout, so it inherits the changed default; its production step
retains a 1h30m limit, and the poller continues sleeping and checking until completion or its own
timeout.

operatorcert/entrypoints/rm_operator_from_index.py[145-151]
operatorcert/iib.py[163-164]
operatorcert/iib.py[189-215]
ansible/roles/operator-pipeline/templates/openshift/tasks/build-fbc-index-images.yml[116-120]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The shared polling default was extended to 210 minutes, but the index-removal caller runs inside a step capped at 90 minutes.

## Fix Focus Areas
- operatorcert/iib.py[163-164]
- operatorcert/entrypoints/rm_operator_from_index.py[145-151]
- ansible/roles/operator-pipeline/templates/openshift/tasks/build-fbc-index-images.yml[116-120]

## Recommended Fix
Give the removal caller an explicit timeout that fits its step if the 90-minute limit is intentional. Otherwise, extend the removal step and its enclosing run limits to accommodate the longer wait.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
⚠️ Tickets: not configured — ticket URL found in PR but could not be fetched — check ticket provider credentials
✅ Compliance rules (platform): 15 rules
✅ Cross-repo context — repo relationships
  Explored: repo: konflux-ci/community-operators-prod (sha: 63702a67) — View relationship
  Explored: repo: release-engineering/iib (sha: 6196e70d) — View relationship
Review mode: Auto: ⚖️ Balanced: Behavioral timeout changes span pipelines and runtime polling logic.

Grey Divider

Tip of the day
💡 Did you know, you can show, collapse, or hide each part of a finding: code, evidence, and all

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread operatorcert/iib.py
@tomasbakk
tomasbakk merged commit e7db98b into main Oct 9, 2026
14 checks passed
@tomasbakk
tomasbakk deleted the ISV-7735 branch October 9, 2026 12:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants