Skip to content

processor: isolate CFL codec state across worker invocations - #12511

Draft
edsiper wants to merge 4 commits into
masterfrom
fix/processor-cfl-worker-state
Draft

edsiper wants to merge 4 commits into
masterfrom
fix/processor-cfl-worker-state

Conversation

@edsiper

@edsiper edsiper commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Consecutive native processors keep a CFL chunk alive after releasing each stage lock. The chunk references the first stage's decoder and encoder, which another output worker can reset while the earlier invocation is still processing. Concurrent mixed chains can lose or mix records or crash. Reproduced on unchanged master 6e438653c5507d32e34364cbd403835903d9eb9c.

Create codecs lazily per flb_processor_run() invocation, reuse them across native segments within that invocation, and claim the actual chunk encoder's buffer before teardown. Clean up the chunk and codecs on every return path. Stage locking stays unchanged; allocation failures call flb_errno() first. No configuration or bundled-library changes. This fixes an existing worker bug separately from performance PR #12505.

CI prerequisite fixes from #12509

The original Linux, ASan and Windows x64 jobs all crashed in
output_exit_destroys_pending_flushes. Locally reproduced with CTest and
Valgrind. The test zeroed task.routes without initializing its list head, so
flb_task_get_route_status() dereferenced an invalid list. Shutdown also retained
processed_event_chunk after freeing it, allowing the flush destructor to free
it again.

Include the two existing commits from #12509 as prerequisites, preserving their
author and DCO trailers: initialize the test route list, and clear the processed
chunk pointer after shutdown cleanup. The output bug remains tracked separately
in #12509; the worker isolation change remains in its own commits.

Verification

Latest CI-fix validation:

  • All 98 internal CTest targets pass, including flb-it-output.
  • flb-it-output passes strict Valgrind: zero errors and zero bytes in use at exit.
  • 15 worker/HTTP integration cases pass normally and strict Valgrind; all 15
    memory logs have zero errors and no definite/indirect/possible leaks.
  • flb-rt-out_http and flb-rt-out_null pass. The HTTP target initially could not
    bind the occupied local port 8888; rerun with FLB_TEST_HTTP_PORT=48995 passed.
  • Full four-commit PR range passes prefix lint against fetched master.
ctest --test-dir build -L internal --output-on-failure
valgrind --error-exitcode=99 --leak-check=full --show-leak-kinds=all \
  --errors-for-leak-kinds=definite,indirect,possible build/bin/flb-it-output
FLB_TEST_HTTP_PORT=48995 ctest --test-dir build \
  -R '^flb-rt-(out_http|out_null)$' --output-on-failure
# With FLUENT_BIT_BINARY pointing at the rebuilt candidate:
python -m pytest tests/integration/scenarios/processor_concurrency \
  tests/integration/scenarios/out_http \
  -k 'worker_threads or sends_json_payload or receiver_error_is_observable or retries_when_chunked_trailer_block_is_invalid' -q
VALGRIND=1 VALGRIND_STRICT=1 python -m pytest \
  tests/integration/scenarios/processor_concurrency tests/integration/scenarios/out_http \
  -k 'worker_threads or sends_json_payload or receiver_error_is_observable or retries_when_chunked_trailer_block_is_invalid' -q

Earlier worker/performance combined verification:

  • Final rebuilt binary: 12 integration cases pass normally and under strict Valgrind Memcheck: input/output processors, ordinary/threaded HTTP inputs, native/mixed/filter-boundary chains, four input workers, four workers per output route, two-route fan-out, eight clients and 4,096 uniquely identified records per route.
  • Three focused CTest targets pass: processor, processor_conditional and mp_chunk_cobj.
  • Combined build with processor: fuse compatible content modifier chains #12505: 40 selected integration cases pass strict Valgrind; normal run passed 38 and hit two metrics-endpoint readiness errors. All three selected readiness reruns passed normally and strict Valgrind after the harness waited for the endpoint.
  • Functional stress tests and Memcheck are not a race-detector proof.
  • Full PR commit range passes the repository prefix linter against fetched master.

From repository root, with FLUENT_BIT_BINARY=$PWD/build/bin/fluent-bit:

python -m pytest tests/integration/scenarios/processor_concurrency -q
VALGRIND=1 VALGRIND_STRICT=1 python -m pytest tests/integration/scenarios/processor_concurrency -q
ctest --test-dir build -R '^flb-it-(processor|processor_conditional|mp_chunk_cobj)$' --output-on-failure

Target master; no backports included.

Signed-off-by: Eduardo Silva <eduardo@chronosphere.io>
Signed-off-by: Eduardo Silva <eduardo@chronosphere.io>
@coderabbitai

coderabbitai Bot commented Oct 2, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Signed-off-by: Hiroshi Hatake <hiroshi@chronosphere.io>
Signed-off-by: Eduardo Silva <eduardo@chronosphere.io>
Signed-off-by: Hiroshi Hatake <hiroshi@chronosphere.io>
Signed-off-by: Eduardo Silva <eduardo@chronosphere.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants