Skip to content

fix(video): stop RTSP tiles hanging on "Connecting…" forever (#77) - #78

Merged
keyldev merged 1 commit into
mainfrom
fix/rtsp-connect-hang
Sep 29, 2026
Merged

keyldev merged 1 commit into
mainfrom
fix/rtsp-connect-hang

Conversation

@keyldev

@keyldev keyldev commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Related

Type

  • Bug fix
  • Feature
  • Refactor / cleanup
  • Docs / CI
  • Other:

Checklist

  • Builds with 0 warnings (TreatWarningsAsErrors=true).
  • Tests pass (dotnet test); new Core logic has unit tests.
  • No layering violation — App references Core only (Infrastructure / Video / Devices wired via DI in a head).
  • Scope stays within one phase (didn't pull work from a later phase's "Не входит").
  • README / docs updated if public commands, options, or setup changed.

Platforms tested

  • Windows
  • Linux
  • macOS
  • Android
  • iOS
  • CI build only

Screenshots / notes

- live sessions passed "stimeout", which FFmpeg 7 ignores, so a camera that
  accepted TCP but never answered blocked avformat_open_input indefinitely
- set the real "timeout" key and wire an interrupt callback with per-call
  budgets (15s open/probe, 10s read, 1s close) plus abort on Dispose
- reconnect watchdog now also abandons attempts stuck in Connecting (45s)
  and uses a monotonic clock
@keyldev
keyldev merged commit a9fc6d1 into main Sep 29, 2026
5 checks passed
@qodo-free-for-open-source-projects

Copy link
Copy Markdown

PR Summary by Qodo

Prevent RTSP live-view tiles from hanging while connecting

🐞 Bug fix 🕐 20-40 Minutes

Grey Divider

AI Description

• Replace the ignored FFmpeg socket option and bound blocking open, probe, read, and close calls.
• Abort blocking I/O on disposal so stalled sessions can shut down.
• Reconnect attempts stuck in Connecting after 45 seconds using a monotonic watchdog.
Diagram

graph TD
  CAM["RTSP camera"] --> FF["FFmpeg session"] --> WD{"Connection stalled?"} --> RE["Reconnect loop"] --> TILE["Tile status"]
  FF --> IO["I/O interrupt"] --> RE
  RE --> FF
  FF --> TILE
Loading
High-Level Assessment

The layered approach fits the failure modes: the corrected socket option handles socket inactivity, per-call interrupts bound longer FFmpeg operations, and the wrapper watchdog catches connection setup outside those calls. A socket timeout or watchdog alone would leave gaps.

Files changed (2) +93 / -17

Bug fix (2) +93 / -17
AutoReconnectingVideoSession.csReconnect sessions stuck in Connecting +35/-14

Reconnect sessions stuck in Connecting

• Adds a 45-second Connecting watchdog alongside the existing five-second frame-stall check. Both checks now use monotonic elapsed time and feed the existing reconnect path.

src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs

FfmpegVideoSession.csBound blocking RTSP operations and correct the socket option +58/-3

Bound blocking RTSP operations and correct the socket option

• Replaces the ignored 'stimeout' option with 'timeout' and installs an interrupt callback before opening the stream. The callback enforces separate open/probe, read, and close budgets and aborts I/O on disposal.

src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs

@keyldev
keyldev deleted the fix/rtsp-connect-hang branch September 29, 2026 19:16
@qodo-free-for-open-source-projects

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (2) 📘 Rule violations (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Wedged connections leave decoder threads alive 🐞 Bug ☼ Reliability
Description
The new Connecting watchdog sends a timed-out attempt through inner.DisposeAsync(), which returns
after a two-second join even if the decoder thread is still blocked in native setup. When that setup
cannot be interrupted, the wrapper starts another attempt while the old thread still holds native
resources and is no longer tracked for disposal.
Code

src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[R259-262]

+                    if (_connecting && idle > ConnectTimeoutMs)
+                    {
+                        failed.TrySetResult("Connect timed out (no stream after 45s)");
+                        return;
Evidence
The PR explicitly targets Connecting operations that the interrupt callback cannot stop, including
hardware decoder setup. The watchdog retries after DisposeAsync, but that method does not require
the decoder thread to exit before returning; native cleanup runs only on that thread.

src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[25-28]
src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[253-263]
src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[198-212]
src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[167-184]
src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[397-410]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The Connecting watchdog retries after disposing an inner session, but disposal can return while its native decoder thread remains blocked. Repeated retries can leave untracked threads and native resources alive.
## Fix Focus Areas
- src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[259-262]
- src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[198-204]
- src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[167-184]
## Recommended Fix
Do not treat a timed-out inner session as fully disposed when its decoder thread has not exited. Retain and track unfinished sessions for eventual cleanup, and bound or defer retries so an uninterruptible setup cannot accumulate decoder threads.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Read timeouts lose their failure reason 🐞 Bug ◔ Observability
Description
Run arms a deadline before av_read_frame, but its negative-result branch logs the returned error
and exits without checking whether that deadline expired or setting LastError. If the read is
interrupted by the new deadline before the frame watchdog signals failure, the wrapper retries with
a null error, so the tile and reconnect log omit the timeout cause.
Code

src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[R308-310]

+                ArmIoDeadline(ReadBudgetMs);
                ret = ffmpeg.av_read_frame(fmtCtx, packet);
                if (ret < 0)
Evidence
The added deadline can interrupt a read, but the existing read-result branch only logs and breaks.
That path reaches Idle without setting an error; the wrapper takes inner.LastError as its retry
reason, and the tile reads that value for its error message.

src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[308-320]
src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[385-414]
src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[416-432]
src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[174-204]
src/OpenIPC.Viewer.App/ViewModels/CameraTileViewModel.cs[329-341]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
An interrupted `av_read_frame` exits as an ordinary idle transition without recording that the new read deadline expired.
## Fix Focus Areas
- src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[308-320]
- src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[416-432]
## Recommended Fix
On a negative read result, distinguish cancellation and EOF from an expired read deadline. Report an expired deadline as a timeout through the session's failure path so `LastError` reaches the reconnect wrapper and UI.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
Review mode: ⚖️ Balanced: Downgraded extended -> standard: change is below the extended eligibility bar (hunks 13/18, lines 110/200; both must reach the floor). Router rationale: This changes native FFmpeg I/O interruption, disposal, timeout budgeting, and reconnect watchdog behavior across two concurrent code paths, creating multiple subtle, high-impact failure modes that benefit from independent review passes.

Grey Divider

Tip of the day
💡 Did you know, you can add REVIEW.md to your repo root and Qodo follows it on every PR

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment on lines +259 to +262
if (_connecting && idle > ConnectTimeoutMs)
{
failed.TrySetResult("Connect timed out (no stream after 45s)");
return;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

1. Wedged connections leave decoder threads alive 🐞 Bug ☼ Reliability

The new Connecting watchdog sends a timed-out attempt through inner.DisposeAsync(), which returns
after a two-second join even if the decoder thread is still blocked in native setup. When that setup
cannot be interrupted, the wrapper starts another attempt while the old thread still holds native
resources and is no longer tracked for disposal.
Agent Prompt
## Issue description
The Connecting watchdog retries after disposing an inner session, but disposal can return while its native decoder thread remains blocked. Repeated retries can leave untracked threads and native resources alive.
## Fix Focus Areas
- src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[259-262]
- src/OpenIPC.Viewer.Video/Pipeline/AutoReconnectingVideoSession.cs[198-204]
- src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[167-184]
## Recommended Fix
Do not treat a timed-out inner session as fully disposed when its decoder thread has not exited. Retain and track unfinished sessions for eventual cleanup, and bound or defer retries so an uninterruptible setup cannot accumulate decoder threads.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +308 to 310
ArmIoDeadline(ReadBudgetMs);
ret = ffmpeg.av_read_frame(fmtCtx, packet);
if (ret < 0)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

2. Read timeouts lose their failure reason 🐞 Bug ◔ Observability

Run arms a deadline before av_read_frame, but its negative-result branch logs the returned error
and exits without checking whether that deadline expired or setting LastError. If the read is
interrupted by the new deadline before the frame watchdog signals failure, the wrapper retries with
a null error, so the tile and reconnect log omit the timeout cause.
Agent Prompt
## Issue description
An interrupted `av_read_frame` exits as an ordinary idle transition without recording that the new read deadline expired.
## Fix Focus Areas
- src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[308-320]
- src/OpenIPC.Viewer.Video/Pipeline/FfmpegVideoSession.cs[416-432]
## Recommended Fix
On a negative read result, distinguish cancellation and EOF from an expired read deadline. Report an expired deadline as a timeout through the session's failure path so `LastError` reaches the reconnect wrapper and UI.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant