Skip to content

f2903ce still enters 100% CPU busy loop during ReleaseCheck #401

Description

@swesty

Summary

pop-upgrade daemon enters a persistent 100%-of-one-core userspace busy loop during a release check on Pop!_OS 22.04, despite running commit f2903ce, which was intended to fix the metadata-fetch infinite retry / 100% CPU issue.

The daemon remained in this state for multiple days. A one-minute watcher recorded 3,791 consecutive samples at approximately 98–101% CPU. APT/DPKG were idle and held no locks.

Environment

  • OS: Pop!_OS 22.04
  • pop-upgrade: 1.0.0~1778861967~22.04~f2903ce (amd64)
  • Installed binary build ID: f549629fb753e75a2c415efd24b915ad2fba5bf4
  • Upstream commit: f2903ce2631c18d721530ee263aff6e31c5951b1
  • Daemon PID in capture: 1233442

Trigger and last daemon messages

The busy loop began after a routine source refresh completed and the daemon entered its release check:

pop-upgrade[1233442]: [INFO ] daemon/mod.rs:1144: updating apt sources
pop-upgrade[...]: Fetched 2798 kB in 2s
pop-upgrade[...]: Reading package lists...
pop-upgrade[1233442]: [INFO ] daemon/mod.rs:1055: performing a release check
pop-upgrade[1233442]: [INFO ] release_api.rs:63: checking for build 24.04 in channel generic

No subsequent completion or error was logged.

The endpoint used by this path was tested separately and returned HTTP 200 with valid release JSON:

https://api.pop-os.org/builds/24.04/generic

Observations

  • CPU stayed at 98–101% of one core for at least 3,791 consecutive one-minute samples.
  • No APT or DPKG lock holder was present.
  • No apt, apt-get, or dpkg child existed beneath pop-upgrade.
  • Unlike pop-upgrade 100% CPU usage #297, there was no defunct apt-get child.
  • strace -f showed the Tokio worker threads waiting in futex/epoll calls, while the CPU-burning main thread made no meaningful system calls.
  • GDB confirmed that thread 1 (the daemon main thread) was running continuously in stripped Rust code. The Tokio worker threads were asleep.
  • The sampled main-thread instruction was in an atomic reference-count cleanup path (lock decq / return), consistent with an async future being polled/dropped repeatedly in userspace.
  • The daemon process mapped the current /usr/bin/pop-upgrade inode; it was not an older, deleted executable left running after package replacement.

Expected behavior

The release API request should either complete successfully or return an error after the configured timeout. The daemon should not spin indefinitely or consume a full CPU core.

Additional evidence available

I have retained:

  • per-minute CPU/thermal/process samples;
  • process tree and /proc state;
  • strace summary and raw trace;
  • six-hour journal capture around detection;
  • complete GDB thread apply all bt output;
  • exact package version, binary build ID, and upstream source revision.

I can attach sanitized captures if useful. The full journal is not attached initially because it may contain unrelated host information.

Possible regression/incomplete fix

The installed revision is the commit from #396, “Remove Infinite Retry on Metadata Fetch To Prevent 100% CPU Usage.” This capture suggests either another busy-loop path remains in the release API request/future handling or the fix does not cover this trigger.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions