BF: restrict value of get_fdmax - #417
Conversation
In my case I kept finding celery running at 100% and doing nothing. py-spy
pointed to the close_open_fds and then ulimit inside the container showed gory
detail of
❯ docker run -it --rm --entrypoint bash dandiarchive/dandiarchive-api -c "ulimit -n"
1073741816
situation is not unique to me. See more at
dandi/dandi-cli#1488
|
hm, why pre-commit.ci is even configured if there is no |
auvipy
left a comment
There was a problem hiding this comment.
lets ignore the pre commit. can you elaborate more on the change please? also should we also consider adding some tests to verify the proposed changes?
|
I would be happy to elaborate! ATM I can only reiterate what tried to describe in original description -- on some systems ulimit would return HUGE number for maximal number of open descriptiors, which would be infeasible to loop through. So, billiard should not try to loop through all the possible billion of them. |
|
May be we can add some unit tests for the suggested changes as well |
There was a problem hiding this comment.
Pull Request Overview
The PR adds logic to cap and warn about excessively high file descriptor limits returned by get_fdmax, preventing performance issues when iterating open descriptors.
- Capture
os.sysconf('SC_OPEN_MAX')intofdmaxand handle errors uniformly - Introduce a threshold (100k) to cap
fdmaxto a sensible default (either the passeddefaultor 10 000) and emit a warning - Import
warningsand emit a deprecation-style warning instead of returning an oversized limit
Comments suppressed due to low confidence (1)
billiard/compat.py:118
- Consider adding unit tests for the new high-value cap branch to verify that large
fdmaxvalues produce the expected warning and capped return value.
if fdmax >= 1e5:
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
|
@yarikoptic @auvipy Thanks for the report and the reproduction — this is the same container setting ( Since #455 landed, What #455 doesn't cover is where |
#456) * List the fd directory in close_open_fds() instead of walking the limit #455 made close_open_fds() call os.closerange() on the gaps between kept descriptors, which is close_range(2) on Linux and finishes in ~0 ms whatever RLIMIT_NOFILE says. Where that syscall is unavailable (seccomp denying it, kernels before 5.9, macOS) os.closerange() falls back to a C loop over the range, still about 3 minutes at the container limit from celery/celery#9886. List /proc/self/fd (or /dev/fd on macOS, and on FreeBSD/DragonFly when fdescfs is mounted) and close only the descriptors actually open, the way CPython's subprocess does; keep the closerange() path as the fallback. get_fdmax() is no longer consulted on the fd directory path, so its value doesn't need capping (cf. #417). Errors from os.close() are ignored, matching os.closerange(). * Handle missing fd directory in test * Patch os.closerange in the remaining fd-directory test Also wrap the skip message from 85456e4 for flake8 and note gh-148575 in the FD_DIR comment, since Cygwin is only there from that change on. --------- Co-authored-by: Asif Saif Uddin {"Auvi":"অভি"} <auvipy@gmail.com>

In my case I kept finding celery running at 100% and doing nothing. py-spy pointed to the close_open_fds and then ulimit inside the container showed gory detail of
situation is not unique to me. See more at
I verified that with this fix my celery container gets unstuck and proceeds to report useful errors ;)