Conversation
This commit forces timeout check of all flows in the flow table at the shutdown stage of Suricata. Gathering of capture-bypassed flow statistics was left to the bypass capture method via BypassUpdate callback. Until now, capture-bypassed flows that did not timeout had their statistics unchecked in the period between last check and shutdown. This commit forces gathering of statistics from these flows. Ticket: 8440
This change prevents capture-bypassed flows to be removed from the flow table by a worker thread. If a flow like this is de-initialized any other way than how FlowManager handles de-initialization of capture-bypassed flows, the underlying bypass method is not aware that it should not filter-out the flow anymore (e.g. EBPF map is not updated). This can lead to a resource leak, such as EBPF map being fully filled out and bypass being incapable of filtering any new flows. Until now, this issue happened in a case when capture-bypassed flow has reached its timeout (as in suricata.yaml flow-timeouts.x.bypassed), but worker thread was the first who pre-emptively de-initialized the flow, before FlowManager could perform proper de-initialization. Ticket: 8442
Decreased default sleep of BypassManager thread from 10ms to 10us, but increased the number of sleep loops adequately to keep the same total sleep time of 10ms. Add configurable option to set number of sleep iterations.
This feature brings capture offload into the DPDK Suricata runmode. The offload is based on the DPDK's rte_flow rules, mainly the drop and count actions. The bypass utilizes the BypassManager thread and API. The number of flows offloaded to the NIC differs based on the used NIC, the current version was tested on only Mellanox ConnectX-6 and supports offload of around 2 millions flows. The real number of rte_flow rules inserted into the NIC is double of the offloaded flows (so around 4 millions rules in this case), because rte_flow needs 2 separate rules to match both directions of a flow. All of the flows that cannot be bypassed in the NIC are offloaded locally in Suricata. The process of applying the bypass is as follows: the workers enqueue a flow_key created from the bypassed flow into a rte_ring, the BypassManager constantly polls the ring, tries to create and insert the respective rte_flow rule into the NIC. The Flow Manager is in the meantime responsible for checking the activity of the flows, performed by reading the counters attached to the rte_flow rules, and potentially destroying the rules in the case the flow timeouts. The statistics from the flows are collected periodically (managed by Flow Manager) and in the shutdown stage and they are captured in eve.json and stats.log. The feature can be toggled on/off and the maximum numbers of capture-bypassed flows (up to NICs maximum) can be set in suricata.yaml. Ticket: 7871
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #15877 +/- ##
==========================================
- Coverage 83.00% 82.82% -0.18%
==========================================
Files 1003 1006 +3
Lines 276582 277261 +679
==========================================
+ Hits 229579 229645 +66
- Misses 47003 47616 +613
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
Contributor
|
Except for the failing QA, does this still need to be a draft? |
Contributor
Author
Actually no, next version can be a regular PR. |
Contributor
Author
|
Continues in #16203 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
DPDK dynamic bypass with rte_flow rules
This is a draft of a feature that brings capture offload into the DPDK Suricata
runmode. The offload is based on the DPDK's rte_flow rules, mainly the drop
and count actions. The bypass utilizes the BypassManager thread and API for
creating bypass rules and the FlowManager for collecting statistics and destroying
rules.
Description
The number of flows offloaded to the NIC differs based on the used NIC,
the current version was tested only on Mellanox ConnectX-6, where it
supports offload of around 2 million flows. The real number of rte_flow
rules inserted into the NIC is double the offloaded flows (maximum around
4 million rules in this case), because rte_flow needs 2 separate rules
to match both directions of a flow. The flows that cannot be
bypassed in the NIC are offloaded in Suricata by software bypass.
The number of rules and the limits of the Mellanox NICs are described in more detail
in this article.
The process of applying the bypass is as follows:
FlowKeystruct created from the bypassed flow into a ring structure,BypassManagerconstantly polls the ring, tries to create and insertthe respective rte_flow rule into the NIC.
FlowManageris, in the meantime, responsible for checking the activity of the flows,reads the counters attached to the bypassed flows, and
potentially destroys the rules in the case of the flow timeouts.
The statistics from the flows are collected periodically and
in the shutdown stage (managed by
FlowManager) and they are captured in eve.json and stats.log.The feature can be toggled on/off and the maximum number of
capture-bypassed flows (up to NICs maximum) can be set in suricata.yaml
Related branches
This PR is based on work in 2 other PRs, which are waiting for their corresponding SV PRs:
Changes:
dpdk.capture-bypasstoggle in suricata.yamlbypass-ring-dequeue-burst-size, which states how many items are dequeued from a ring in one BypassManager Iterationbypass-info-mp-size, which configures how many flows can be rte_flow bypass handle (up to a NIC specific maximum)flow.bypass.delay, which defines how long BypassManager sleeps in one iteration. The value is stated in hundreds of microsecondsBenchmarks
The goal of this benchmark is to compare the flow bypass performance of Suricata's main version with that of the newly implemented hardware-based rte_flow bypass. Both cases include running Suricata in the DPDK runmode.
All configuration files and PCAPs are included in the ticket description.
Hardware setup and specification:
Our setup includes 2 servers: one Suricata server and one server running TRex traffic generator. The traffic flows between 2 TRex network interfaces. One of the interfaces on TRex is mirrored to Suricata's interface; this way, we get both Rx and Tx traffic.
Suricata machine:
rte_flow bypass parameters
The main 3 parameters that can be tweaked are
flow.bypass.delay,dpdk.bypass-ring-sizeanddpdk.bypass-ring-dequeue-burst-sizeflow.bypass.delay: This value stands for how many hundreds of microseconds the BypassManager thread will sleep in its main loop, e.g., how much time the thread will sleep between calls to the subprogram where the hardware rules are created. For example, setting the value to 5 will lead to 500 microseconds of sleep in one pass of the main loop.bypass-ring-size: This parameter configures how many flows can wait at the same time for their corresponding HW rule to be created. The flows that do not fit inside thebypass-ringwill be picked up by the software bypass.bypass-ring-dequeue-burst-size: This number states how many rules will be created in one iteration of BypassManager's main loop, e.g., how many items will be dequeued from thebypass-ringIn the benchmarks, we are maximizing the throughput while aiming at 1% drop-rate. We are testing 1, 10 and 100 values for
flow.bypass.delay, each with a different combination ofbypass-ring-sizeandbypass-ring-dequeue-burst-sizeSmall flows benchmark
In this test, we are sending a shorter flow towards Suricata, which uses ETO rule-set + one bypass rule, that bypasses each flow after the first packet. The included TRex profile is tuned to run at 10Gbps with a multiplier set to 1.
sudo ./t-rex-64 -f /etc/https.yaml -d 10 -m x -c 10flow.bypass.delaybypass-ring-sizebypass-ring-dequeue-burst-sizeResults
The base version of Suricata achieved ~1% drop-rate at approximately 27Gbps in this testing setup. As visible in the table above, the best results were achieved using configurations with a large ring and optimally with smaller BypassManager sleeps. In the best scenarios, we have achieved a 40% increase in throughput.
Large flows benchmark
In this test, we are running Suricata with ETO rule-set and we are testing bypass on larger flows (~12MB), where we trigger the bypass after 1MB of the flow has been analyzed. This means that around 11 out of 12 parts of the flow should avoid inspection. The included TRex profile is tuned to run at 10Gbps with multiplier set to 1.
sudo ./t-rex-64 -f /etc/big_flows.yaml -d 30 -m x -c 10flow.bypass.delaybypass-ring-sizebypass-ring-dequeue-burst-sizeResults
When generating larger flows, the base Suricata DPDK version achieved ~1% drop-rate at around 17Gbps, while the Suricata with rte_flow rules achieved this drop-rate at 34Gbps. This marks a 100% higher throughput with the rte_flow bypass enabled.
As seen from the results, adjusting the
flow.bypass.delaydid not bring any changes to the final measured throughput. Thebypass-ring-sizeandbypass-ring-dequeue-burst-sizealso have no effect on throughput.First packet captured Benchmark
This test measures how many packets are passed through local software bypass before the rte_flow hardware rule takes effect. The flows are bypassed from the first packet.
The test was conducted by generating a PCAP with 1 flow of 100k packets and replaying it at 1M pps. Each resulting value is the average of 5 distinct measurements. The table contains the configuration of
flow.bypass.delay, minimum, maximum, and average number of packets that were inspected by Suricata before the bypass took effect.We tested it using TRex in stateless mode, replaying the traffic from the TRex mirrored interface. We have set the IPG (stated in microseconds) to 1.
flow.bypass.delayAs visible from the table above, having a more frequent polling mechanism leads to a quicker reaction, the rules are created earlier and this results in fewer packets being inspected directly by Suricata.
Link to ticket: #7871
Previous PR: #15290