Share TCP has carried most network traffic for decades, but modern data centers place very different demands on transport protocols. Microservices generate thousands of short RPCs, distributed systems can create sudden incast traffic, and AI and HPC clusters require high bandwidth with very low latency and minimal CPU overhead. This is where Homa, DCTCP, RDMA and RoCE enter the discussion. While standard TCP provides a reliable, ordered byte stream, DCTCP uses ECN-based congestion feedback to keep data center switch queues under control. Homa takes an RPC-oriented approach, using receiver-driven scheduling and message priorities to reduce tail latency for short requests. RDMA and RoCE address another bottleneck by allowing network adapters to move data with far less CPU and kernel involvement. This comparison explains TCP vs Homa vs DCTCP vs RDMA/RoCE, including congestion control, incast, receiver scheduling, switch queues, kernel bypass, Azure Accelerated Networking, and RDMA networking for AI and HPC workloads. TCP vs HOMA TCP vs HOMA Four questions Compare Incast DCTCP Homa RDMA Matrix Simulator Azure Flowchart Tree Sequence Radar Stacks FAQ Data-center transport explained TCP vsHOMA TCP moves a reliable stream of bytes. Homa moves RPC messages and lets the receiver decide who sends next. Here is why that difference matters for tail latency, and where DCTCP and RDMA fit in. Research notes · Indrajit Ghosh · October 2026 TCP · SENDERS PUSH HOMA · RECEIVER GRANTS queue builds near-empty grant → uncoordinated sendsgranted, shortest first 01 / THE SHORT VERSION Four protocols, four different questions These technologies do not all solve the same problem. Each one asks a different question about how data should move. TCP“How much data can I send without overwhelming the network?” DCTCP“How much congestion is actually happening, so I can cut my sending rate by the right amount?” HOMA“Which RPC should the receiver let in next, so short messages don’t get stuck behind long ones?” RDMA“Why involve the CPU and kernel at all? Let the NIC move data directly between memories.” RDMA is not another TCP competitor. TCP, DCTCP, and Homa define transport behavior. RDMA defines a memory-access model, and RoCEv2, iWARP, and InfiniBand are ways to carry RDMA traffic. 02 / HEAD TO HEAD TCP vs Homa at a glance TCP is a general-purpose, connection-oriented byte-stream protocol. Homa is an RPC-oriented transport built for very low latency inside data centers. AreaTCPHoma Primary designGeneral Internet transportData-center RPC transport Communication modelContinuous byte streamDiscrete messages / RPCs ConnectionConnection-orientedConnectionless, message-oriented SetupHandshake before dataNo traditional connection setup OrderingIn-order byte deliveryDelivers complete RPC messages Congestion controlSender-driven (CUBIC, BBR, DCTCP variants)Receiver-driven scheduling Short RPC latencyCan suffer from queues and head-of-line blockingBuilt for short-message tail latency Large transfersVery mature and effectiveSupported, but designed around RPC workloads MultiplexingUsually needs multiple connectionsMany concurrent RPCs on one socket DeploymentUniversalResearch / specialized CompatibilitySupported everywhereNeeds Homa-capable hosts and stack The biggest difference: who controls the send button In TCP, the sender decides when to transmit, limited only by its congestion window. When many senders target one server, they all decide to send at once, and queues build. In Homa, a sender can push a small first slice of an RPC without asking. After that, the receiver hands out grants that decide which sender goes next. TCP: sender decides Application ↓ Byte stream ↓ TCP connection ↓ Congestion window ↓ Network Homa: receiver decides Application RPC ↓ Message ↓ Initial unscheduled packets ↓ Receiver grants transmission ↓ Priority scheduling ↓ Network A worked example A server receives three RPCs at the same time: A = 2 KB, B = 20 KB, and C = 5 MB. With TCP, packets from all three compete in the same queues, so the tiny RPC can sit behind packets from the 5 MB transfer. Homa puts the short ones first: TCP: shared queue Queue: [C][C][B][C][A][C][C][B][C]... A waits behind C's packets Homa: shortest first 2 KB → 20 KB → rest of 5 MB A ✓ B ✓ C ... That is why Homa focuses on tail latency (p99, p99.9) rather than average throughput. Try it yourself in the simulator. 03 / THE PROBLEM Why TCP struggles inside a data center TCP sees bytes, not messages. If an application sends RPC A (2 KB), RPC B (5 MB), and RPC C (4 KB), TCP sees one long run of bytes from 0 to about 5,006,000. It has no idea that C is tiny and latency-sensitive while B is huge. In microservice architectures, that blind spot is expensive. Incast: many senders, one receiver A front-end service fans a request out to 100 backends. Each replies with 64 KB at roughly the same moment: 100 × 64 KB = 6.4 MB converging on one receiver. The senders together can push hundreds of gigabits, but the receiver link may be only 100 Gbps. Packets pile up in the top-of-rack or leaf switch. Fan-out, then fan-in Request │ ┌────────▼────────┐ │ Front-end API │ └────────┬────────┘ fan out to 100 servers ┌─────────────┼─────────────┐ ▼ ▼ ▼ Server 1 Server 2 ... Server 100 └─────────────┼─────────────┘ ▼ 6.4 MB lands at once What the switch sees Servers │ │ │ │ │ │ │ │ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼ ┌──────────────────┐ │ Switch queue │ │ ███████████████ │ ← growing └────────┬─────────┘ 100 Gbps ▼ Receiver The chain reaction: queueing → latency → buffer exhaustion → drops → retransmissions → even more latency. This is how p99 and p99.9 latency get ugly while the average still looks healthy. 04 / KEEP TCP, FIX THE FEEDBACK DCTCP: earlier, gentler warnings Microsoft Research built Data Center TCP around this exact problem. The team measured a 6,000-server production cluster and found that bandwidth-heavy background flows were building queues in switches and hurting latency-sensitive foreground traffic. DCTCP keeps everything familiar about TCP: connections, sequence numbers, ACKs, byte-stream semantics, and congestion windows. What changes is the feedback. Instead of waiting for a full buffer and a dropped packet, switches use ECN (Explicit Congestion Notification) to mark packets once the queue passes a threshold. The sender then measures how much of its traffic was marked and cuts its window in proportion, instead of treating congestion as a yes/no event. Traditional TCP Queue ███████████████████████ X packet loss ↓ slow down hard DCTCP Queue ██████ ↑ ECN mark at threshold K ↓ slow down in proportion 90%less switch buffer space, with the same or better throughput than TCPDCTCP, SIGCOMM 2010 10×more background traffic handled without hurting foreground trafficDCTCP, SIGCOMM 2010 2021SIGCOMM Test of Time Award for the DCTCP paperMicrosoft Research The limit: DCTCP is still reactive. Congestion starts, the network marks it, and the sender responds. It shrinks queues; it does not schedule them. 05 / A DIFFERENT STARTING POINT Homa: schedule the receiver’s link Homa starts from a simple observation: most data-center traffic is RPCs, so why model it as byte streams? Instead of connection → stream → bytes, Homa works with RPC → message → packets. Because it knows where each message starts and ends, it knows how much of each one is left to send. Shortest Remaining Processing Time (SRPT) Homa approximates SRPT. Rather than splitting bandwidth equally, it gets short messages out of the way first. Equal sharing RPC A ████ RPC B ████ RPC C ████ RPC D ████ (everyone progresses slowly) SRPT 2 KB ████ ✓ 4 KB ██████ ✓ 20 KB ████████████ ✓ 5 MB ████████████████████... Unscheduled and scheduled packets A sender may transmit the first part of a message right away, without permission. These are unscheduled packets, and they let short messages finish in one shot. The rest of a larger message moves as scheduled packets, sent only when the receiver issues a grant. TCP: everyone sends Sender A ──┐ Sender B ──┤ Sender C ──┤ Sender D ──┼──► Receiver Sender E ──┤ Sender F ──┘ Potential incast. Homa: receiver grants Sender A ─────┐ Sender B ─────┤ Sender C ─────┼──── Receiver Sender D ─────┘ │ remaining bytes: A: 2 KB B: 5 MB C: 8 KB ┌────────┼─────────┐ ▼ ▼ ▼ grant A grant C wait B Why receiver control works A receiver with a 100 Gbps NIC can only ever take in 100 Gbps. Under TCP, 50 machines might each try to send toward it at line rate, and the network has to absorb the mismatch. Homa lets the receiver answer one question: who gets my downlink next? The bottleneck is scheduled, not just reacted to. 06 / A DIFFERENT BOTTLENECK RDMA, RoCE, and InfiniBand Homa and DCTCP deal with network congestion. RDMA deals with CPU overhead. At 100, 200, or 400+ Gbps, pushing millions of packets through the kernel stack burns a lot of CPU. Normal TCP receive path NIC ↓ interrupt / polling ↓ kernel network stack ↓ TCP processing ↓ socket buffer ↓ copy ↓ application RDMA: kernel bypass, zero-copy Application memory │ ▼ NIC │ =======NETWORK======= │ NIC │ ▼ Remote application memory The NIC places data straight into registered application memory with very little CPU involvement. The result is very low latency, very high throughput, and low CPU use. That is why RDMA dominates distributed AI training, MPI/HPC, high-performance storage, and GPU-to-GPU communication. RoCEv2: RDMA over Ethernet RoCEv2 stack Application │ RDMA verbs │ RDMA NIC │ RoCEv2 │ UDP / IP │ Ethernet What the fabric needs PFC Priority Flow Control ETS Enhanced Transmission Selection ECN Explicit Congestion Notification Goal: make packet loss extremely rare RDMA expects a network where loss almost never happens, and Ethernet does not promise that on its own. Microsoft’s Azure Local guidance calls for careful DCB configuration, including PFC and ETS, when using RoCE. Getting a RoCE fabric right takes real network engineering work. Two attacks on two bottlenecks Simplified, but useful Network congestion ▲ │ Homa attacks this │ TCP ────────────────────┼──────────── │ RDMA attacks this │ ▼ CPU overhead The two ideas can coexist. The Homa Linux research found that once network congestion was under control, software processing overhead became the main limit, and the authors argued that future gains would need transport processing moved into NIC hardware. 07 / THE MATRIX How they compare Stars are conceptual ratings, not benchmark numbers. CharacteristicTCPDCTCPHomaRDMA / RoCE Internet compatibility★★★★★★✕✕ General-purpose★★★★★★★★★★★★★ Short RPC latency★★★★★★★★★★★★★★★ Tail latency★★★★★★★★★★★★★★★ CPU efficiency★★★★★★★★★★★★★★ Throughput★★★★★★★★★★★★★★★★★★ Handles incast★★★★★★★★★★★Depends on the fabric Deployment simplicity★★★★★★★★★★★ Commodity ecosystem★★★★★★★★★★★★★ RPC awareness✕✕✓Application dependent Homa’s benchmark, in context 7 to 83×lower p99 latency for short messages than TCP and DCTCP, across tested workloadsHoma Linux, USENIX ATC 2021 40nodes in the test clusterHoma Linux, USENIX ATC 2021 Read this carefully: this does not mean “Homa is 83× faster than TCP.” It is a workload-specific research result that shows what happens when transport is redesigned around data-center RPCs. Homa is not a mainstream transport, and research on receiver-driven designs continues. For example, the 2025 SIRD work (NSDI) compares itself against Homa, DCTCP, Swift, and dcPIM. 08 / TRY IT Queue simulator Three RPCs arrive at one receiver at the same moment. Compare a FIFO switch queue with fair sharing (TCP-like) against priority scheduling with shortest-first ordering (Homa-like). Illustrative model, not a benchmark bars use a log scale RPC A size RPC B size RPC C size Queued backlog at switch Receiver link Fixed overhead per RPC Default Heavy backlog Empty queue How it works. Link rate in bytes per µs = Gbps × 125. TCP-like: every RPC first waits for the backlog to drain (FIFO), then the three RPCs share the link equally until each finishes. Homa-like: short RPCs use higher priority queues, so they skip the backlog and go in size order; the largest RPC shares the lowest priority with background traffic. Both add the same fixed overhead (software, propagation). 1 KB = 1,000 bytes. Real results depend on workload, topology, and implementation. 09 / AZURE Where this lands in Azure Different Azure scenarios sit in different parts of this picture. That distinction is easy to miss. Regular cloud application App Service ↓ AKS ↓ Redis ↓ Azure SQL = TCP/IP world + Accelerated Networking HPC / AI training GPU │ matrix operation │ AllReduce │ network ← stall here = idle GPUs │ GPU │ next operation = RDMA over InfiniBand Ordinary Azure applications: TCP/IP For web apps, APIs, and databases, you are in TCP/IP. The levers are connection reuse, HTTP/2 or HTTP/3, good placement (for example proximity placement groups), Private Link, and right-sized VMs and NICs. Accelerated Networking enables SR-IOV so traffic goes from the NIC straight to the VM, bypassing the host and virtual switch while network policies are enforced in hardware. This lowers latency, jitter, and CPU use. For best results, enable it on at least two VMs in the same virtual network; it has little effect on latency across virtual networks or to on-premises. For most Azure architectures, this matters far more than Homa. Azure HPC and AI: RDMA over InfiniBand Most Azure HPC VM sizes, plus selected N-series sizes marked with “r”, add a second network interface for RDMA over InfiniBand alongside the standard Ethernet NIC. Almost all newer RDMA-capable sizes are SR-IOV enabled. RDMA runs only over the InfiniBand network, not over Ethernet. To use RDMA between VMs, place them in the same virtual machine scale set or availability set, and keep your virtual network address space clear of 172.16.0.0/16, which the RDMA network reserves. Azure Local: RoCE and iWARP On Azure Local (on-premises, Azure-connected infrastructure), Microsoft supports RDMA through both RoCE and iWARP. RoCE needs the DCB settings described above. Quick decision guide TCP / QUICInternet-facing apps and APIsWeb traffic, file transfer, databases across WANs. Use Accelerated Networking on Azure VMs. DCTCPControlled data center, TCP appsYou own the switches and hosts, want smaller queues, and need to keep TCP semantics. HOMAHyperscale RPC researchMillions of short RPCs between machines where tens of microseconds at p99 matter. Not a production Azure option. RDMAAI training, HPC, storageInfiniBand on Azure HPC and N-series “r” VMs. RoCE or iWARP on Azure Local. Template: explaining this to a partner Copy-ready summary Copy Subject: TCP, Homa, and RDMA: what applies to your Azure workload Hi [Name], Quick summary of the transport options we discussed: 1. Your [web/API/database] tier runs on TCP/IP. The practical gains come from Accelerated Networking (SR-IOV), connection reuse, and keeping chatty VMs close together in the same virtual network. 2. Homa is a research transport that schedules RPCs at the receiver and favors short messages. It shows large p99 improvements in lab tests, but it is not a mainstream or Azure-supported option today. 3. For [AI training / HPC / MPI] workloads, the right path is RDMA over InfiniBand using RDMA-capable Azure HPC or N-series “r” VM sizes, deployed in the same scale set or availability set. 4. For Azure Local, RDMA is supported through RoCE or iWARP. RoCE needs DCB (PFC and ETS) configured on the switches. Next step: [share VM sizes / workload profile] so we can confirm the right configuration. Thanks, [Your name] 10 / DECISION FLOWCHART Which transport fits your workload? Answer the questions on the right. The flowchart lights up the path you take and ends at the recommendation from the decision guide above. Interactive decision path Start over 11 / DECOMPOSITION TREE Break the problem down, branch by branch Every technology on this page attacks one of two bottlenecks. Click any node to expand or collapse it. Data-center transport, decomposed Expand all Collapse Reset Network queues CPU and data copies Azure mapping Has hidden children 12 / SEQUENCE DIAGRAM A Homa exchange, step by step Two senders target one receiver: A has a 2 KB RPC, C has a 5 MB RPC. Press play to watch unscheduled packets, grants, and scheduled packets in order. Receiver-driven scheduling ▶ Play ◀ Back Next ▶ 0Press play or Next to start. 13 / RADAR VIEW The matrix as a shape The same star ratings from section 07, plotted on eight axes. Toggle each protocol to compare profiles. Conceptual profile 0 to 5 stars · ✕ = 0 Handles incast and RPC awareness are left out because RDMA’s values there are “depends on the fabric” and “application dependent”. 14 / LAYER STACKS Where each one changes the stack Gradient layers show what each technology replaces or adds. Striped layers are skipped. Data path, top to bottom TCP Applicationbyte stream Kernel stack + socket buffercopy TCPsender congestion window NIC Networkloss = signal DCTCP Applicationbyte stream Kernel stack + socket buffer DCTCPwindow cut in proportion to marks NIC NetworkECN marks at threshold K Homa Application RPCmessages Kernel stacksoftware overhead is the new limit Homareceiver grants · SRPT · priorities NIC Networksmall queues RDMA (RoCEv2) Application memoryRDMA verbs Kernel stack + copy RDMA NICtransport offloaded RoCEv2 · UDP/IP EthernetPFC · ETS · ECN 15 / FAQ Frequently asked questions Is Homa a replacement for TCP?No. Homa targets RPC traffic inside a low-latency data center. For Internet traffic, WAN database connections, and general file transfer, TCP or QUIC remains the practical choice. Can I use Homa on Azure today?Not as a supported Azure networking feature. Homa needs Homa-capable hosts and a Homa network stack on both ends. On Azure, latency work for regular apps centers on Accelerated Networking and placement; for HPC and AI it centers on RDMA over InfiniBand. What is the difference between DCTCP and Homa?DCTCP keeps TCP and improves congestion feedback using ECN, so senders slow down in proportion to congestion. Homa replaces the byte stream with messages and lets the receiver schedule which sender transmits next, favoring the shortest remaining message. DCTCP reacts to queues; Homa tries to prevent them. Does Homa’s 7 to 83× result mean it is always that much faster?No. Those numbers are p99 latency for short messages on a 40-node test cluster across specific workloads. They show the potential of RPC-aware transport, not a general speedup. Why does RDMA appear next to transport protocols if it is not one?Because it competes for the same job: moving data between machines fast. RDMA attacks CPU and memory-copy overhead, while Homa and DCTCP attack network queueing. They solve different bottlenecks and can coexist. What is incast?Many senders replying to one receiver at the same time. The combined traffic exceeds the receiver’s link, the switch queue fills, packets drop, and retransmissions push tail latency up. Does Accelerated Networking give me RDMA?No. Accelerated Networking uses SR-IOV to send Ethernet traffic directly to the VM, which cuts latency, jitter, and CPU use for TCP/IP workloads. RDMA on Azure runs over the separate InfiniBand network on RDMA-capable VM sizes. Why is RoCE harder to run than InfiniBand?RDMA expects almost no packet loss. InfiniBand is built for that. Ethernet is not, so RoCE needs extra configuration such as PFC, ETS, and ECN to keep loss rare. 16 / SOURCES Sources Data Center TCP (DCTCP), Microsoft Research, SIGCOMM 2010 A Linux Kernel Implementation of the Homa Transport Protocol, USENIX ATC 2021 (PDF) SIRD, USENIX NSDI 2025 Azure Accelerated Networking overview, Microsoft Learn Set up InfiniBand on HPC VMs, Microsoft Learn Azure Local host network requirements, Microsoft Learn Azure Local physical network requirements, Microsoft Learn Disclaimer: The Questions and Answers provided on https://gigxp.com are for general information purposes only. We make no representations or warranties of any kind, express or implied, about the completeness, accuracy, reliability, suitability or availability with respect to the website or the information, products, services, or related graphics contained on the website for any purpose. Share What's your reaction? Excited 0 Happy 0 In Love 0 Not Sure 0 Silly 0 IG Website Twitter
Big Data Polyglot Persistence: Strategic Guide to Modern Data Architecture In today’s complex application landscape, the one-size-fits-all database is a relic. Modern systems, from e-commerce ...
Cloud Computing Comparing Cosmos DB vs MongoDB vs DynamoDB NoSQL DBs Choosing the right NoSQL database is one of the most critical decisions for any modern, ...
Cloud Computing DirectLake vs Athena vs Redshift Spectrum: The Ultimate 2025 Lakehouse BI Guide In the complex world of data analytics, the 2025 lakehouse landscape is dominated by two ...
Cloud Computing Power BI on AWS: The Ultimate Guide to Athena, Glue & Lake Formation with RLS Connecting Power BI to a modern AWS data lake is a powerful strategy for enterprise ...
Cloud Computing Power BI on Mac: 10 Best Alternatives for 2025 Are you a Mac user struggling to fit into a Power BI-centric world? The frustration ...
Big Data Build your Data Estate with Azure DataBricks – Part 3 – IoT It is not the strongest of the species that survives, nor the most intelligent that ...
Cloud Computing Google Sheets Docs vs Office 365 – Excel Online Word Suite Comparison If you are looking to switch to the cloud, you have two options at your ...
Cloud Computing Webex Cellular Data Consumption in Approximately 1 Hour + Estimation Online business collaboration tools are in vogue these days and have been the most important ...
Cloud Computing What Windows Server roles are not Supported on Azure Virtual machines? If you are into software or IT infrastructure, you should be aware that Azure Virtual ...
Cloud Computing Differences Between VMWare SRM Standard and Enterprise Licensing Disaster Management is one of the essential aspects of any virtualization technique. VMWare, being one ...
Cloud Computing Google Cloud vs Alibaba Cloud Provider Feature Comparisons and Pricing Cloud services have been one of the prominent options in the current times. Among the ...
Cloud Computing Cisco Webex Bandwidth Requirements – Network QoS Internet & Data Use Cloud-based collaboration tools are one of the best options for increased productivity. The suite for ...