Skip to content

Skip Route lookups during cleanup on clusters without the Route API - #2

Merged
lwr20 merged 1 commit into
rel/v3.24.0-2from
fix/route-cleanup-recursionerror-non-openshift
Sep 1, 2026
Merged

lwr20 merged 1 commit into
rel/v3.24.0-2from
fix/route-cleanup-recursionerror-non-openshift

Conversation

@lwr20

@lwr20 lwr20 commented Sep 1, 2026 •

Copy link
Copy Markdown
Member

Problem

On a plain-Kubernetes cluster the ForkliftController never reaches Successful. It settles at:

Successful=False   Failure=True (Failed)   Running=False
message: MODULE FAILURE: No start of json char found   (x3)
ansibleResult: {"changed":0,"failures":0,"ok":62,"skipped":62}

The install itself is fine — 6/6 deployments Available, 6/6 cert-manager Certificates and both Issuers Ready=True. Only the status is wrong.

Cause

forklift_resources (defaults/main.yml:102) is [Deployment, ConfigMap, Service, Route], and the four disabled-feature cleanup loops in tasks/main.yml iterate it unguarded. The enable paths are all correctly gated on not k8s_cluster|bool (lines 325/370/393); the cleanup loops are not.

So the documented plain-Kubernetes CR — which sets feature_ui_plugin: "false", feature_cli_download: "false" — drives cleanup for those features, and cleanup.yml calls k8s_info with kind: Route. With no route.openshift.io on the cluster, the dynamic client recurses:

File ".../kubernetes/dynamic/resource.py", in __search
    return self.__search(parts[1:], resourcePart, reqParams + [part])
RecursionError: maximum recursion depth exceeded

The module writes a traceback to stderr and nothing to stdout, so Ansible reports the opaque MODULE FAILURE: No start of json char found, naming neither Route nor the recursion.

The precise trigger is an unqualified kind — cleanup.yml passes kind with no api_version, so the dynamic client searches every group and __search recurses without terminating when nothing matches. Verified in the operator image against this cluster:

k8s_info kind=Route                                              -> FAILED, RecursionError
k8s_info api_version=console.openshift.io/v1 kind=ConsolePlugin  -> SUCCESS, "Failed to find API for resource..."

A kind qualified with its api_version fails gracefully. This is worth knowing for two reasons: the finalizer's Remove console plugin is not affected (its template names apiVersion: console.openshift.io/v1), and the fix below addresses the current resource list rather than the underlying class — any future OpenShift-only kind added to forklift_resources would reintroduce it. Carrying api_version alongside each kind would be the structural fix, but the four kinds span different groups, so that is a larger change than this release needs.

cleanup.yml already has a rescue: for "empty or missing resources" and it works — hence failures: 0 and a healthy install. What it cannot suppress is the operator recording the module-failure events, which is what pins the CR at Failure=True.

Confirmed deterministic: three failing tasks per reconcile (Get Route resources labeled forklift-{ui-plugin,cli-download,mcp-server}), 13 distinct Ansible job IDs over ~3h, identical every time.

Impact

  • The documented install step kubectl wait --for=condition=Successful forkliftcontroller/... --timeout=300s always times out.
  • Any automation gating on Successful hangs.
  • Failure/Successful stop being usable health signals on Kubernetes — a genuine later failure is indistinguishable from this noise.

Fix

Filter Route out of the cleanup kinds when the API group is absent, reusing the api_groups detection the role already does two lines above. Keying off api_groups rather than k8s_cluster means the decision follows what the cluster can actually resolve, not a user-settable flag.

Testing

RKE2 v1.33.5, 5 nodes, Calico Enterprise v3.24.0-2.0-calient-1.dev-151, Forklift v2.12.5-v3.24.0-2.0, installed from the published non-OLM manifests.

  • Reproduced before the change, as above.
  • Jinja verified both ways: with route.openshift.io present the list is unchanged; absent, it is [Deployment, ConfigMap, Service].
  • A full vSphere -> Kubernetes migration succeeds on this cluster either way (the install is functional; only the CR status is affected).

Note

cleanup.yml and the forklift_resources list are byte-identical to kubev2v/forklift@main, so this affects every non-OpenShift Forklift user and is worth raising upstream too. This PR fixes it on the release branch.

The disabled-feature cleanup loops iterate forklift_resources, which includes
Route. On a cluster with no route.openshift.io, k8s_info recurses in the dynamic
client and dies with RecursionError. Ansible reports only "MODULE FAILURE: No
start of json char found", and although cleanup.yml rescues the task, the
operator still records the module failure and pins the ForkliftController at
Failure=True.

The install itself completes, so the effect is that the CR never reaches
Successful on Kubernetes: the documented "kubectl wait --for=condition=Successful"
always times out, and Failure/Successful stop being usable health signals there.

Filter Route out of the cleanup kinds when the API group is absent, using the
same api_groups detection the role already performs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 1, 2026 16:16

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This pull request fixes a reconcile-time failure on plain Kubernetes clusters by preventing the disabled-feature cleanup tasks from querying OpenShift-only Route resources when the route.openshift.io API group is not available, avoiding Ansible k8s_info module failures that incorrectly pin the CR status to Failure=True.

Changes:

  • Introduces a derived cleanup_resources list based on detected cluster API groups, filtering out Route when route.openshift.io is absent.
  • Updates all disabled-feature cleanup loops to iterate cleanup_resources instead of forklift_resources.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@lwr20
lwr20 merged commit a334509 into rel/v3.24.0-2 Sep 1, 2026
2 checks passed
lwr20 added a commit that referenced this pull request Sep 1, 2026
The disabled-feature cleanup loops iterate forklift_resources, which includes
Route. cleanup.yml passes the kind to k8s_info with no api_version, so on a
cluster without route.openshift.io the dynamic client searches every group and
__search recurses until it dies with RecursionError. Ansible surfaces only
"MODULE FAILURE: No start of json char found", and although cleanup.yml rescues
the task, the operator still records the module failure and pins the
ForkliftController at Failure=True.

The install itself completes, so the effect is that the CR never reaches
Successful on Kubernetes: the documented "kubectl wait --for=condition=Successful"
always times out, and Failure/Successful stop being usable health signals there.

Filter Route out of the cleanup kinds when the API group is absent, using the
same api_groups detection the role already performs.

Cherry-pick of a334509 from rel/v3.24.0-2 (#2).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lwr20 added a commit that referenced this pull request Sep 1, 2026
…n-openshift-main

Skip Route lookups during cleanup on clusters without the Route API (cherry-pick of #2)
aaaaaaaalex pushed a commit that referenced this pull request Sep 15, 2026
The disabled-feature cleanup loops iterate forklift_resources, which includes
Route. cleanup.yml passes the kind to k8s_info with no api_version, so on a
cluster without route.openshift.io the dynamic client searches every group and
__search recurses until it dies with RecursionError. Ansible surfaces only
"MODULE FAILURE: No start of json char found", and although cleanup.yml rescues
the task, the operator still records the module failure and pins the
ForkliftController at Failure=True.

The install itself completes, so the effect is that the CR never reaches
Successful on Kubernetes: the documented "kubectl wait --for=condition=Successful"
always times out, and Failure/Successful stop being usable health signals there.

Filter Route out of the cleanup kinds when the API group is absent, using the
same api_groups detection the role already performs.

Cherry-picked from #2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alex O'Regan <alex.oregan@tigera.io>
aaaaaaaalex pushed a commit that referenced this pull request Sep 16, 2026
The disabled-feature cleanup loops iterate forklift_resources, which includes
Route. cleanup.yml passes the kind to k8s_info with no api_version, so on a
cluster without route.openshift.io the dynamic client searches every group and
__search recurses until it dies with RecursionError. Ansible surfaces only
"MODULE FAILURE: No start of json char found", and although cleanup.yml rescues
the task, the operator still records the module failure and pins the
ForkliftController at Failure=True.

The install itself completes, so the effect is that the CR never reaches
Successful on Kubernetes: the documented "kubectl wait --for=condition=Successful"
always times out, and Failure/Successful stop being usable health signals there.

Filter Route out of the cleanup kinds when the API group is absent, using the
same api_groups detection the role already performs.

Cherry-picked from #2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Alex O'Regan <alex.oregan@tigera.io>
mrnold pushed a commit to kubev2v/forklift that referenced this pull request Sep 16, 2026
…8566)

The disabled-feature cleanup loops iterate forklift_resources, which
includes Route. cleanup.yml passes the kind to k8s_info with no
api_version, so on a cluster without route.openshift.io the dynamic
client searches every group and __search recurses until it dies with
RecursionError. Ansible surfaces only "MODULE FAILURE: No start of json
char found", and although cleanup.yml rescues the task, the operator
still records the module failure and pins the ForkliftController at
Failure=True.

The install itself completes, so the effect is that the CR never reaches
Successful on Kubernetes: "kubectl wait --for=condition=Successful"
always times out, and Failure/Successful stop being usable health
signals there.

Filter Route out of the cleanup kinds when the API group is absent,
using the same api_groups detection the role already performs.

Cherry-picked from tigera#2.

Signed-off-by: Alex O'Regan <alex.oregan@tigera.io>
Co-authored-by: Lance Robson <lance@tigera.io>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants