Skip to content

DPDK rte_flow dynamic bypass draft v4 - #15877

Closed
adaki4 wants to merge 8 commits into
OISF:mainfrom
adaki4:dpdk-rte-flow-dynamic-bypass-draft-v4
Closed

adaki4 wants to merge 8 commits into
OISF:mainfrom
adaki4:dpdk-rte-flow-dynamic-bypass-draft-v4

Conversation

@adaki4

@adaki4 adaki4 commented Jul 19, 2026 •

Copy link
Copy Markdown
Contributor

DPDK dynamic bypass with rte_flow rules

This is a draft of a feature that brings capture offload into the DPDK Suricata
runmode. The offload is based on the DPDK's rte_flow rules, mainly the drop
and count actions. The bypass utilizes the BypassManager thread and API for
creating bypass rules and the FlowManager for collecting statistics and destroying
rules.

Description

The number of flows offloaded to the NIC differs based on the used NIC,
the current version was tested only on Mellanox ConnectX-6, where it
supports offload of around 2 million flows. The real number of rte_flow
rules inserted into the NIC is double the offloaded flows (maximum around
4 million rules in this case), because rte_flow needs 2 separate rules
to match both directions of a flow. The flows that cannot be
bypassed in the NIC are offloaded in Suricata by software bypass.
The number of rules and the limits of the Mellanox NICs are described in more detail
in this article.

The process of applying the bypass is as follows:

  1. Workers enqueue a FlowKey struct created from the bypassed flow into a ring structure,
  2. BypassManager constantly polls the ring, tries to create and insert
    the respective rte_flow rule into the NIC.
  3. FlowManager is, in the meantime, responsible for checking the activity of the flows,
    reads the counters attached to the bypassed flows, and
    potentially destroys the rules in the case of the flow timeouts.

The statistics from the flows are collected periodically and
in the shutdown stage (managed by
FlowManager) and they are captured in eve.json and stats.log.

The feature can be toggled on/off and the maximum number of
capture-bypassed flows (up to NICs maximum) can be set in suricata.yaml

Related branches

This PR is based on work in 2 other PRs, which are waiting for their corresponding SV PRs:

Changes:

  • fixed dpdk.capture-bypass toggle in suricata.yaml
  • added more bypass statistics and their descriptions
  • added option bypass-ring-dequeue-burst-size, which states how many items are dequeued from a ring in one BypassManager Iteration
  • added option bypass-info-mp-size, which configures how many flows can be rte_flow bypass handle (up to a NIC specific maximum)
  • added option flow.bypass.delay, which defines how long BypassManager sleeps in one iteration. The value is stated in hundreds of microseconds
  • performed benchmarks compared to the base Suricata
  • rebased to main

Benchmarks

The goal of this benchmark is to compare the flow bypass performance of Suricata's main version with that of the newly implemented hardware-based rte_flow bypass. Both cases include running Suricata in the DPDK runmode.

All configuration files and PCAPs are included in the ticket description.

Hardware setup and specification:

Our setup includes 2 servers: one Suricata server and one server running TRex traffic generator. The traffic flows between 2 TRex network interfaces. One of the interfaces on TRex is mirrored to Suricata's interface; this way, we get both Rx and Tx traffic.

Suricata machine:

  • CPU: Intel(R) Xeon(R) Silver 4114 CPU @ 2.20GHz
  • CPU cores: Used 18 CPU cores on 1 NUMA node
  • NIC: MT2892 Family [ConnectX-6 Dx]
  • Memory: 64 GB
  • Hugepages: 18 GB

rte_flow bypass parameters

The main 3 parameters that can be tweaked are flow.bypass.delay, dpdk.bypass-ring-size and dpdk.bypass-ring-dequeue-burst-size

flow.bypass.delay: This value stands for how many hundreds of microseconds the BypassManager thread will sleep in its main loop, e.g., how much time the thread will sleep between calls to the subprogram where the hardware rules are created. For example, setting the value to 5 will lead to 500 microseconds of sleep in one pass of the main loop.

bypass-ring-size: This parameter configures how many flows can wait at the same time for their corresponding HW rule to be created. The flows that do not fit inside the bypass-ring will be picked up by the software bypass.

bypass-ring-dequeue-burst-size: This number states how many rules will be created in one iteration of BypassManager's main loop, e.g., how many items will be dequeued from the bypass-ring

In the benchmarks, we are maximizing the throughput while aiming at 1% drop-rate. We are testing 1, 10 and 100 values for flow.bypass.delay, each with a different combination of bypass-ring-size and bypass-ring-dequeue-burst-size

Small flows benchmark

In this test, we are sending a shorter flow towards Suricata, which uses ETO rule-set + one bypass rule, that bypasses each flow after the first packet. The included TRex profile is tuned to run at 10Gbps with a multiplier set to 1.

  • PCAP: https.pcap, single flow, 174 packets, 174163 bytes
  • Traffic duration: 10 seconds
  • suricata.yaml: suricata-bypass-benchmark.yaml
  • Command: sudo ./t-rex-64 -f /etc/https.yaml -d 10 -m x -c 10
Configuration flow.bypass.delay bypass-ring-size bypass-ring-dequeue-burst-size ring.occupancy_max ring.occupancy_avg Throughput
Base case — — — — — 27Gbps
Small Ring, Small Burst 1 256 64 255 6 35Gbps
Medium Ring, Small Burst 1 8192 64 8191 322 38Gbps
Medium Ring, Medium Burst 1 8192 1024 8191 21 38Gbps
Large Ring, Large Burst 1 16384 16384 15125 41 38Gbps
Small Ring, Small Burst 10 256 64 255 62 34Gbps
Medium Ring, Small Burst 10 8192 64 8191 2626 33Gbps
Medium Ring, Medium Burst 10 8192 1024 8191 81 34Gbps
Large Ring, Large Burst 10 16384 16384 13792 119 36Gbps
Small Ring, Small Burst 100 256 64 255 231 29Gbps
Medium Ring, Small Burst 100 8192 64 8191 7227 28Gbps
Medium Ring, Medium Burst 100 8192 1024 8191 1319 35Gbps
Large Ring, Large Burst 100 16384 16384 16383 987 37Gbps
Very Large Ring, Very Large Burst 100 32768 32768 16800 864 37Gbps
image

Results

The base version of Suricata achieved ~1% drop-rate at approximately 27Gbps in this testing setup. As visible in the table above, the best results were achieved using configurations with a large ring and optimally with smaller BypassManager sleeps. In the best scenarios, we have achieved a 40% increase in throughput.

Large flows benchmark

In this test, we are running Suricata with ETO rule-set and we are testing bypass on larger flows (~12MB), where we trigger the bypass after 1MB of the flow has been analyzed. This means that around 11 out of 12 parts of the flow should avoid inspection. The included TRex profile is tuned to run at 10Gbps with multiplier set to 1.

  • PCAP: netflix.pcap, single flow, 11002 packets, 12382524 bytes
  • Traffic Duration: 30 seconds
  • suricata.yaml: suricata-bypass-benchmark.yaml
  • Command: sudo ./t-rex-64 -f /etc/big_flows.yaml -d 30 -m x -c 10
Configuration flow.bypass.delay bypass-ring-size bypass-ring-dequeue-burst-size ring.occupancy_max ring.occupancy_avg Throughput
Base case — — — — — 17Gbps
Small Ring, Small Burst 1 256 64 8 1 34Gbps
Small Ring, Small Burst 10 256 64 7 1 34Gbps
Small Ring, Small Burst 100 256 64 12 4 34Gbps
Large Ring, Large Burst 1 16384 16384 6 1 34Gbps
Large Ring, Large Burst 10 16384 16384 4 1 34Gbps
Large Ring, Large Burst 100 16384 16384 11 4 34Gbps
image

Results

When generating larger flows, the base Suricata DPDK version achieved ~1% drop-rate at around 17Gbps, while the Suricata with rte_flow rules achieved this drop-rate at 34Gbps. This marks a 100% higher throughput with the rte_flow bypass enabled.

As seen from the results, adjusting the flow.bypass.delay did not bring any changes to the final measured throughput. The bypass-ring-size and bypass-ring-dequeue-burst-size also have no effect on throughput.

First packet captured Benchmark

This test measures how many packets are passed through local software bypass before the rte_flow hardware rule takes effect. The flows are bypassed from the first packet.

The test was conducted by generating a PCAP with 1 flow of 100k packets and replaying it at 1M pps. Each resulting value is the average of 5 distinct measurements. The table contains the configuration of flow.bypass.delay, minimum, maximum, and average number of packets that were inspected by Suricata before the bypass took effect.

We tested it using TRex in stateless mode, replaying the traffic from the TRex mirrored interface. We have set the IPG (stated in microseconds) to 1.

flow.bypass.delay MIN MAX AVG
1 911 1032 972
10 1608 2287 1892
100 2879 12575 8225

As visible from the table above, having a more frequent polling mechanism leads to a quicker reaction, the rules are created earlier and this results in fewer packets being inspected directly by Suricata.

image

Link to ticket: #7871

Previous PR: #15290

adaki4 added 8 commits July 17, 2026 17:43
This commit forces timeout check of all flows in the flow table at the
shutdown stage of Suricata.

Gathering of capture-bypassed flow statistics was left to the bypass
capture method via BypassUpdate callback. Until now, capture-bypassed
flows that did not timeout had their statistics unchecked in the period
between last check and shutdown. This commit forces gathering of
statistics from these flows.

Ticket: 8440
This change prevents capture-bypassed flows to be removed from the
flow table by a worker thread. If a flow like this is de-initialized
any other way than how FlowManager handles de-initialization
of capture-bypassed flows, the underlying bypass method is not aware
that it should not filter-out the flow anymore (e.g. EBPF map is not
updated). This can lead to a resource leak, such as EBPF map being
fully filled out and bypass being incapable of filtering any new flows.

Until now, this issue happened in a case when capture-bypassed flow has
reached its timeout (as in suricata.yaml flow-timeouts.x.bypassed),
but worker thread was the first who pre-emptively de-initialized
the flow, before FlowManager could perform proper de-initialization.

Ticket: 8442
Decreased default sleep of BypassManager thread from 10ms to 10us, but
increased the number of sleep loops adequately to keep the same total
sleep time of 10ms.

Add configurable option to set number of sleep iterations.
This feature brings capture offload into the DPDK Suricata runmode.
The offload is based on the DPDK's rte_flow rules, mainly the drop
and count actions. The bypass utilizes the BypassManager thread and API.

The number of flows offloaded to the NIC differs based on the used NIC,
the current version was tested on only Mellanox ConnectX-6 and
supports offload of around 2 millions flows. The real number of rte_flow
rules inserted into the NIC is double of the offloaded flows (so around
4 millions rules in this case), because rte_flow needs 2 separate rules
to match both directions of a flow. All of the flows that cannot be
bypassed in the NIC are offloaded locally in Suricata.

The process of applying the bypass is as follows: the workers enqueue  a
flow_key created from the bypassed flow into a rte_ring,
the BypassManager constantly polls the ring, tries to create and insert
the respective rte_flow rule into the NIC. The Flow Manager is in the
meantime responsible for checking the activity of the flows,
performed by reading the counters attached to the rte_flow rules, and
potentially destroying the rules in the case the flow timeouts.

The statistics from the flows are collected periodically (managed by
Flow Manager) and in the shutdown stage and they are captured in
eve.json and stats.log.

The feature can be toggled on/off and the maximum numbers of
capture-bypassed flows (up to NICs maximum) can be set in suricata.yaml.

Ticket: 7871
@codecov

codecov Bot commented Jul 19, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 18.46785% with 596 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.82%. Comparing base (8455efd) to head (9abef0c).

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #15877      +/-   ##
==========================================
- Coverage   83.00%   82.82%   -0.18%     
==========================================
  Files        1003     1006       +3     
  Lines      276582   277261     +679     
==========================================
+ Hits       229579   229645      +66     
- Misses      47003    47616     +613     
Flag Coverage Δ
fuzzcorpus 61.52% <0.00%> (+0.02%) ⬆️
livemode 18.44% <18.49%> (-0.01%) ⬇️
netns 22.89% <75.00%> (-0.02%) ⬇️
pcap 45.36% <51.72%> (-0.06%) ⬇️
suricata-verify 67.00% <55.17%> (-0.05%) ⬇️
unittests 58.46% <4.16%> (-0.01%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@lukashino

Copy link
Copy Markdown
Contributor

Except for the failing QA, does this still need to be a draft?

@adaki4

adaki4 commented Aug 19, 2026 •

Copy link
Copy Markdown
Contributor Author

Except for the failing QA, does this still need to be a draft?

Actually no, next version can be a regular PR.

@adaki4

adaki4 commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

Continues in #16203

@adaki4 adaki4 closed this Sep 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants