Disclaimer: this post and the code it describes were written with AI assistance (Claude Opus).
Hardware: BlueField-3 (verified) · BlueField-2 (verified with an earlier revision) · ConnectX-7 (TBC)
TL;DR mlx5 has no flow director (FDIR) for IP protocols it does not parse. On BlueField-3, DPDK can already do it through the flex parser, so Linux was only missing the driver interface. We added it to the mlx5 driver. The NIC can now steer such a protocol to RX queues by its own port fields, and traffic between one pair of hosts spreads over many queues and cores.
Background
Anyone who knows me knows I have been working with Homa for a while. Homa is a datacenter transport protocol (SIGCOMM 2018 paper, Linux implementation at USENIX ATC 2021), and its current Linux implementation runs directly on top of IP under its own protocol number, 146, the way TCP and UDP do. Like TCP, its header starts with a source port and a destination port.
A modern NIC has many RX queues, and it uses RSS to decide which queue each packet goes to. The kernel’s scaling document describes the mechanism, and Cloudflare’s post on receiving a million packets per second walks through it on real hardware. The NIC hashes some header fields and uses the hash to pick a queue. For TCP and UDP the hash covers the addresses and the ports. For any other protocol number the NIC can only hash the IP addresses.
For Homa this means that all packets between one pair of hosts get the same hash, land on the same RX queue, and are processed by the same core, however many flows run between those two hosts. That core does all the receive work and becomes the bottleneck.
Same IP pair, same RX queue
To see the effect without Homa, we used protocol 253, one of the two numbers reserved for experiments by RFC 3692. A Python raw-socket sender sent 160,000 packets to a BlueField-3, at about 300,000 packets per second, with a Homa-like header whose destination port cycles over 16 values.
All 160,000 packets went to a single RX queue. In three runs, 205 to 3,356 of them were dropped and counted as rx_out_of_buffer, because the ring of that one queue ran out of buffers. The other 31 queues stayed idle.
The usual workarounds each have a cost:
- More IP addresses per host. The Cloudflare post uses multiple receive IPs to spread UDP traffic. It spreads the load only across address pairs, and the extra addresses end up in every application’s configuration.
- RPS (scaling document) spreads the work in software, after one core has already taken the interrupt and touched every packet.
- Disguising the protocol as TCP or UDP on the wire gets RSS, but middleboxes and the host stack then see the wrong protocol.
Flow steering on other NICs
I wanted the NIC to steer the packets by the protocol’s own port fields. Intel NICs have supported this for years. Their Flow Director (ixgbe, i40e, ice) lets an ethtool rule match a two-byte pattern at an offset into the packet, through the user-def field, which is enough to steer by a port. Cloudflare’s low-latency post uses Flow Director on an Intel 82599 to pin flows to queues.
NVIDIA’s mlx5 driver has no equivalent. Switching NICs was not an option either. Homa relies on some hardware features of NVIDIA NICs, so we could not test it on Intel ones.
The flex parser on NVIDIA NICs
On some NVIDIA NICs the hardware can do this. ConnectX NICs parse packets with a fixed parser that knows Ethernet, VLAN, IP, TCP, UDP and a list of tunnels. The firmware of some NVIDIA NICs also has a programmable flex parser. A new header is described as a node in a parse graph: which existing node it follows and on what condition, how long it is, and which of its dwords to sample. The firmware returns a sample id, and flow-steering rules can match on that sample. We have used it on BlueField-2 and BlueField-3. A standalone ConnectX-6 Dx does not offer it, and we have not tested ConnectX-7.
DPDK exposes the flex parser as the mlx5 “flex item”, which needs two persistent firmware settings, FLEX_PARSER_PROFILE_ENABLE=4 and PROG_PARSE_GRAPH=1. We took the object layout from DPDK. The Linux driver uses flex parser resources only for a few fixed protocols: GENEVE TLV options, GTP-U, MPLS over UDP or GRE, ICMP and VXLAN-GPE. From Linux there was no way to use it for any other protocol.
Connecting it to ethtool
For the user interface we used a rule type that ethtool already has. Besides rules for TCP and UDP, ethtool has ip4 and ip6 rules for any IP protocol, and these rules can match the first four bytes after the IP header. The man page calls these bytes l4data, and the kernel calls them l4_4_bytes (include/uapi/linux/ethtool.h). For a header that starts with two 16-bit ports, l4data holds both ports. The stock mlx5 driver rejects any rule that sets this field (en_fs_ethtool.c), in the 6.18 LTS kernel and in mainline and net-next as of October 2026:
receiver$ ethtool -N ens1f1np1 flow-type ip4 l4proto 253 l4data 3 m 0xfffffff0 action 3
rmgr: Cannot insert RX class rule: Invalid argument
We changed the driver to accept these rules and to implement them with the flex parser. On the first l4data rule, the driver adds one node to the parse graph:
Ethernet ──► IP ─┬─ proto 6 ──► TCP fixed parser, RSS hashes the ports
├─ proto 17 ──► UDP fixed parser, RSS hashes the ports
└─ proto N ──► new node: 4-byte header, sample dword 0
│
ethtool rule: (sample & mask) == l4data ──► RX queue
The node is created with one firmware command:
MLX5_SET(general_obj_in_cmd_hdr, hdr, opcode, MLX5_CMD_OP_CREATE_GENERAL_OBJECT);
MLX5_SET(general_obj_in_cmd_hdr, hdr, obj_type, MLX5_OBJ_TYPE_FLEX_PARSE_GRAPH);
/* A 4-byte header, sampled as dword 0 at a fixed offset */
MLX5_SET(parse_graph_flex, flex, header_length_mode, MLX5_GRAPH_NODE_LEN_FIXED);
MLX5_SET(parse_graph_flex, flex, header_length_base_value, 4);
MLX5_SET(parse_graph_flow_match_sample, sample, flow_match_sample_en, 1);
MLX5_SET(parse_graph_arc, arc, arc_parse_graph_node, MLX5_GRAPH_ARC_NODE_IP);
MLX5_SET(parse_graph_arc, arc, compare_condition_value, proto);
err = mlx5_cmd_exec(mdev, in, sizeof(in), out, sizeof(out));
The driver then queries the node for the firmware-assigned sample id. Each rule matches the sample in the flow-table entry, together with the usual address and protocol match:
MLX5_SET(fte_match_set_misc4, misc4, prog_sample_field_id_0, sample_id);
MLX5_SET(fte_match_set_misc4, misc4, prog_sample_field_value_0, be32_to_cpu(val & mask));
The rules hold references to the node, and deleting the last rule destroys it. The driver manages its GENEVE TLV option object the same way (lib/geneve.c).
The node is attached to the generic IP node, which compares the IPv4 protocol and the IPv6 next header in the same way. One node serves both families, and IPv6 support only needs the driver to accept flow-type ip6 rules.
Protocols the NIC already parses
The firmware also creates a node for a protocol that the fixed parser already handles, and the node then replaces the fixed parser’s handling of it. With a node on GRE (protocol 47), the inner-header RSS of all GRE traffic collapsed to one queue: 64,000 GRE packets that normally spread over 24 queues landed on one, and 7,016 of them were dropped. Deleting the node restored the spread.
So the driver rejects l4data for the protocols the NIC parses: TCP, UDP, ICMP, IPIP, IPv6-in-IP, GRE and ESP. The node is shared by both IP families, so it also rejects ICMPv6 and the IPv6 extension headers, including for IPv4 rules.
Spreading traffic with rules
The RSS hash cannot take a flex parser sample as input. The NIC’s hash input selector only accepts the fixed parser’s TCP and UDP ports, and the DPDK documentation states that flex item fields “do not take part in RSS hash functions”. The traffic has to be spread with exact-match rules. Sixteen rules on the low four bits of the destination port give sixteen buckets:
receiver$ ethtool -K ens1f1np1 ntuple on
receiver$ for k in $(seq 0 15); do
ethtool -N ens1f1np1 flow-type ip4 l4proto 253 l4data $k m 0xfffffff0 action $k
done
Added rule with ID 1023
Added rule with ID 1022
...
l4data is the first four header bytes read as a big-endian number, so for a header that starts with two 16-bit ports it is (sport << 16) | dport. m is ethtool’s mask of bits to ignore, so 0xfffffff0 compares only the low nibble of the destination port. The same rules with flow-type ip6 cover IPv6.
With the rules installed, the same 160,000 packets from the sender spread over the first 16 queues. These are the per-queue deltas of ethtool -S:
receiver$ ./tools/rxq.sh ens1f1np1 rxq.before
rx0_packets 10000
rx1_packets 10000
rx2_packets 10000
rx3_packets 10001
...
rx15_packets 10000
Each queue received the 10,000 packets of its destination port. The extra packet on queue 3 came from other traffic on the interface.
Because each rule names its queue, the placement is predictable. A server that listens on ports P to P+15, with one thread per port pinned to cores 0 to 15 and the interrupt of queue j pinned to core j, can send port P+j to queue j. A packet is then received and consumed on the same core.
Results
The receiver was a BlueField-3 (integrated ConnectX-7, firmware 32.47.3576) running Linux 6.18.55 with the patched driver, and the sender was the one described above. Each configuration ran three times over IPv4:
| RX queues used | busiest queue | delivered | dropped (rx_out_of_buffer) | |
|---|---|---|---|---|
| no rules (stock behaviour) | 1 | all of them | 156,644 to 159,795 | 205 to 3,356 |
16 l4data rules | 16 | 10,000 | 160,000 | 0 |
The rules over IPv6 split 160,000 packets the same way. Other checks on this machine:
- One exact (sport, dport) rule delivers exactly its 1,000 packets to the chosen queue, over IPv4 and over IPv6, and the rest of the traffic takes the default path.
- IPv4 and IPv6 rules for one protocol share the node, and a rule for a second protocol is rejected.
- While the node exists, UDP traffic keeps its RSS spread over all 32 queues,
udp4ntuple rules work, and IPv6 ping is unaffected. - 50 add/delete cycles across four protocols all succeed, with no leak.
- Removing the module with rules installed and reloading it works.
On a ConnectX-4 Lx, which has no programmable parser, the driver checks the capability and rejects the rule:
receiver$ ethtool -N enp23s0f0np0 flow-type ip4 l4proto 253 l4data 0 m 0xfffffff0 action 1
rmgr: Cannot insert RX class rule: Operation not supported
Trying it
The code, a build script and the two tools used above are at github.com/breakertt/mlx5-flex-l4-steering. Requirements:
- a NIC whose firmware has the programmable parser. BlueField-3 works, and an earlier revision worked on BlueField-2. ConnectX-7 may work; we have not tested it. A standalone ConnectX-6 Dx does not, because its firmware does not offer the setting.
- the two persistent firmware settings, set with
mstconfigand activated by a firmware reset. Back up the NV configuration first. - the patched
mlx5_core, built in a kernel tree or out of tree with the script.
Limitations:
- One network interface per PCI function, which is the normal case. Setups with several interfaces on one function, such as SR-IOV switchdev representors, are not handled.
- One such protocol per port at a time. A rule for a second protocol is rejected with
EBUSYuntil the first protocol’s rules are deleted. - Each bucket takes one flow-table rule.
Final words
The flex parser on BlueField-3 can steer an IP protocol that the NIC does not parse, using that protocol’s own header fields. The mlx5 Linux driver lacked the interface, and we connected ethtool’s existing l4data field to it. For Homa, a pair of hosts no longer has to funnel all receive work into one queue and one core.
Thanks to Boris for the hint that got this started.
References
- Homa: SIGCOMM 2018 paper, USENIX ATC 2021 paper on the Linux implementation, HomaModule source, upstream patch series on netdev
- Linux kernel: Scaling in the Linux Networking Stack (RSS, RPS, RFS)
- Cloudflare: How to receive a million packets per second
- Cloudflare: How to achieve low latency with 10Gbps Ethernet
- Cloudflare: How to drop 10 million packets per second (ethtool ntuple filters for dropping in hardware)
- Intel Flow Director: ixgbe, i40e and ice driver documentation
- ethtool(8) man page
- Linux v6.18.55 source:
include/uapi/linux/ethtool.h,mlx5/core/en_fs_ethtool.c,mlx5/core/lib/geneve.c - DPDK: NVIDIA mlx5 poll mode driver, flex item
- RFC 3692: Assigning Experimental and Testing Numbers, IANA protocol numbers
- This work: github.com/breakertt/mlx5-flex-l4-steering