Repository navigation
Tight container limits may cause "read init-p: connection reset by peer" #1914
Description
Activity
- added 2 commits that reference this issue
on Oct 30, 2018 Hi @danail-branekov,
I'm trying to get a better understanding of the series of actions that lead to this error.
- The process
/bin/echo hiis created but hasn't started execution when this line is executed - The process from step (1) is placed into the limited cgroup
- The process from step (1) begins execution.
- Error occurs. The golang runtime fails to create a goroutine/system thread to execute step (3). A goroutine other than the main goroutine is created because
cmd.Start()doesn't wait for the process to exit.
Is this happening because goroutines are assigned their own pids?
- The process
danail-branekov commented
on Nov 12, 2018 ContributorAuthorMore actionsHi @kkallday
Kind of. Thepids/pidmaxcgroup file, despite its name, not only limits the number of allowed process identifiers (PIDs) but also limits the number of thread identifiers (TIDs). When a process starts a new thread (such as a go routine) via e.g.pthread_create, its thread id (TID) adds up to thepidscgroup as well, hence the cgroup limit is reached.As noted above, if you artificially wait some time for the golang runtime go routines finish their initialisation and complete, then the issue is gone as the user process
/bin/echo hiconsists of a single PID/TID.Reacted by jianhaiqing@danail-branekov gotcha - I didn't know that TIDs count towards the
pidscgroup count. That is good to know.One more question: when you
sleepfor some time, does that mean the command might execute for some time outside of the cgroup? After the process is placed in the cgroup, the process might have already exited which defeats the purpose of putting it in a cgroup (I'm assuming my understanding here is wrong). Or is there some type of "freeze" on the command before it gets executed in the cgroup.I'm new here - trying to get a better understanding of the project. Thanks in advance! 😄
danail-branekov commented
on Nov 13, 2018 ContributorAuthorMore actionsWell, yes, sleeping is just a hack/workaround to prove that we get the
connection reseterror because of hitting the pids limit. AFAIK, there is no process freezing, it just starts and its pid is added to the cgroup. Therefore there is some tiny theoretical interval (from here to here) where the process can terminate in the meanwhile. I am not sure what would the behaviour be in that case, maybe running the process would fail...Ideally we would know when we should contain the process with a liveliness check of the Go runtime but that's not really doable (you could try to do it by doing even more synchronisation -- but that's what #1916 will do implicitly).
- added a commit that references this issue
on Oct 8, 2026
Steps to reproduce:
runc create <id>echo 1 > /sys/fs/cgroup/pids/.../<id>/pids.maxrunc exec <id> /bin/echo hi. The following error occurs:After some debugging we found out what causes this error:
In order to prove that we added a sleep of 100ms before the process is joined to the cgroup and this significantly reduced the failure rate. Removing the code that joins the cgroup "fixed" it entirely.
We realise that such a tight limit has quite a limited practical use but we wanted to share the knowledge with the community. We believe that this error may also occur when exceeding any container cgroup limit (e.g. memory, cpu, pids).
Cheers, CF Garden Team