Skip Route lookups during cleanup on clusters without the Route API - #2
Merged
lwr20 merged 1 commit intoSep 1, 2026
Merged
Conversation
The disabled-feature cleanup loops iterate forklift_resources, which includes Route. On a cluster with no route.openshift.io, k8s_info recurses in the dynamic client and dies with RecursionError. Ansible reports only "MODULE FAILURE: No start of json char found", and although cleanup.yml rescues the task, the operator still records the module failure and pins the ForkliftController at Failure=True. The install itself completes, so the effect is that the CR never reaches Successful on Kubernetes: the documented "kubectl wait --for=condition=Successful" always times out, and Failure/Successful stop being usable health signals there. Filter Route out of the cleanup kinds when the API group is absent, using the same api_groups detection the role already performs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
This pull request fixes a reconcile-time failure on plain Kubernetes clusters by preventing the disabled-feature cleanup tasks from querying OpenShift-only Route resources when the route.openshift.io API group is not available, avoiding Ansible k8s_info module failures that incorrectly pin the CR status to Failure=True.
Changes:
- Introduces a derived
cleanup_resourceslist based on detected cluster API groups, filtering outRoutewhenroute.openshift.iois absent. - Updates all disabled-feature cleanup loops to iterate
cleanup_resourcesinstead offorklift_resources.
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
lwr20
added a commit
that referenced
this pull request
Sep 1, 2026
The disabled-feature cleanup loops iterate forklift_resources, which includes Route. cleanup.yml passes the kind to k8s_info with no api_version, so on a cluster without route.openshift.io the dynamic client searches every group and __search recurses until it dies with RecursionError. Ansible surfaces only "MODULE FAILURE: No start of json char found", and although cleanup.yml rescues the task, the operator still records the module failure and pins the ForkliftController at Failure=True. The install itself completes, so the effect is that the CR never reaches Successful on Kubernetes: the documented "kubectl wait --for=condition=Successful" always times out, and Failure/Successful stop being usable health signals there. Filter Route out of the cleanup kinds when the API group is absent, using the same api_groups detection the role already performs. Cherry-pick of a334509 from rel/v3.24.0-2 (#2). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lwr20
added a commit
that referenced
this pull request
Sep 1, 2026
…n-openshift-main Skip Route lookups during cleanup on clusters without the Route API (cherry-pick of #2)
aaaaaaaalex
pushed a commit
that referenced
this pull request
Sep 15, 2026
The disabled-feature cleanup loops iterate forklift_resources, which includes Route. cleanup.yml passes the kind to k8s_info with no api_version, so on a cluster without route.openshift.io the dynamic client searches every group and __search recurses until it dies with RecursionError. Ansible surfaces only "MODULE FAILURE: No start of json char found", and although cleanup.yml rescues the task, the operator still records the module failure and pins the ForkliftController at Failure=True. The install itself completes, so the effect is that the CR never reaches Successful on Kubernetes: the documented "kubectl wait --for=condition=Successful" always times out, and Failure/Successful stop being usable health signals there. Filter Route out of the cleanup kinds when the API group is absent, using the same api_groups detection the role already performs. Cherry-picked from #2. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Alex O'Regan <alex.oregan@tigera.io>
aaaaaaaalex
pushed a commit
that referenced
this pull request
Sep 16, 2026
The disabled-feature cleanup loops iterate forklift_resources, which includes Route. cleanup.yml passes the kind to k8s_info with no api_version, so on a cluster without route.openshift.io the dynamic client searches every group and __search recurses until it dies with RecursionError. Ansible surfaces only "MODULE FAILURE: No start of json char found", and although cleanup.yml rescues the task, the operator still records the module failure and pins the ForkliftController at Failure=True. The install itself completes, so the effect is that the CR never reaches Successful on Kubernetes: the documented "kubectl wait --for=condition=Successful" always times out, and Failure/Successful stop being usable health signals there. Filter Route out of the cleanup kinds when the API group is absent, using the same api_groups detection the role already performs. Cherry-picked from #2. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Alex O'Regan <alex.oregan@tigera.io>
mrnold
pushed a commit
to kubev2v/forklift
that referenced
this pull request
Sep 16, 2026
…8566) The disabled-feature cleanup loops iterate forklift_resources, which includes Route. cleanup.yml passes the kind to k8s_info with no api_version, so on a cluster without route.openshift.io the dynamic client searches every group and __search recurses until it dies with RecursionError. Ansible surfaces only "MODULE FAILURE: No start of json char found", and although cleanup.yml rescues the task, the operator still records the module failure and pins the ForkliftController at Failure=True. The install itself completes, so the effect is that the CR never reaches Successful on Kubernetes: "kubectl wait --for=condition=Successful" always times out, and Failure/Successful stop being usable health signals there. Filter Route out of the cleanup kinds when the API group is absent, using the same api_groups detection the role already performs. Cherry-picked from tigera#2. Signed-off-by: Alex O'Regan <alex.oregan@tigera.io> Co-authored-by: Lance Robson <lance@tigera.io> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On a plain-Kubernetes cluster the
ForkliftControllernever reachesSuccessful. It settles at:The install itself is fine — 6/6 deployments Available, 6/6 cert-manager Certificates and both Issuers
Ready=True. Only the status is wrong.Cause
forklift_resources(defaults/main.yml:102) is[Deployment, ConfigMap, Service, Route], and the four disabled-feature cleanup loops intasks/main.ymliterate it unguarded. The enable paths are all correctly gated onnot k8s_cluster|bool(lines 325/370/393); the cleanup loops are not.So the documented plain-Kubernetes CR — which sets
feature_ui_plugin: "false",feature_cli_download: "false"— drives cleanup for those features, andcleanup.ymlcallsk8s_infowithkind: Route. With noroute.openshift.ioon the cluster, the dynamic client recurses:The module writes a traceback to stderr and nothing to stdout, so Ansible reports the opaque
MODULE FAILURE: No start of json char found, naming neither Route nor the recursion.The precise trigger is an unqualified kind —
cleanup.ymlpasseskindwith noapi_version, so the dynamic client searches every group and__searchrecurses without terminating when nothing matches. Verified in the operator image against this cluster:A kind qualified with its
api_versionfails gracefully. This is worth knowing for two reasons: the finalizer'sRemove console pluginis not affected (its template namesapiVersion: console.openshift.io/v1), and the fix below addresses the current resource list rather than the underlying class — any future OpenShift-only kind added toforklift_resourceswould reintroduce it. Carryingapi_versionalongside each kind would be the structural fix, but the four kinds span different groups, so that is a larger change than this release needs.cleanup.ymlalready has arescue:for "empty or missing resources" and it works — hencefailures: 0and a healthy install. What it cannot suppress is the operator recording the module-failure events, which is what pins the CR atFailure=True.Confirmed deterministic: three failing tasks per reconcile (
Get Route resources labeled forklift-{ui-plugin,cli-download,mcp-server}), 13 distinct Ansible job IDs over ~3h, identical every time.Impact
kubectl wait --for=condition=Successful forkliftcontroller/... --timeout=300salways times out.Successfulhangs.Failure/Successfulstop being usable health signals on Kubernetes — a genuine later failure is indistinguishable from this noise.Fix
Filter
Routeout of the cleanup kinds when the API group is absent, reusing theapi_groupsdetection the role already does two lines above. Keying offapi_groupsrather thank8s_clustermeans the decision follows what the cluster can actually resolve, not a user-settable flag.Testing
RKE2 v1.33.5, 5 nodes, Calico Enterprise
v3.24.0-2.0-calient-1.dev-151, Forkliftv2.12.5-v3.24.0-2.0, installed from the published non-OLM manifests.route.openshift.iopresent the list is unchanged; absent, it is[Deployment, ConfigMap, Service].Note
cleanup.ymland theforklift_resourceslist are byte-identical tokubev2v/forklift@main, so this affects every non-OpenShift Forklift user and is worth raising upstream too. This PR fixes it on the release branch.