TCP, UDP, and Transport Services

A progressive guide to transport-layer communication, explaining ports, TCP reliability and control mechanisms, UDP datagrams, and how to choose between the two protocols.

Transport-Layer Responsibilities

Transport protocols provide process-to-process communication between applications running on different hosts. They operate above IP: IP delivers packets between host addresses, while the transport layer identifies the destination application and may add reliability, ordering, , or .

The two foundational Internet choices are and . offers a feature-rich byte stream, while offers lightweight independent datagrams. The correct choice depends on the communication semantics an application needs, not simply on whether it is described as fast or reliable.

Takeaway: The transport layer converts host-to-host packet delivery into communication between application endpoints, with and offering different service models.

Ports and Application Endpoints

A identifies an application endpoint on a host. and carry source and destination ports, allowing many applications to share one IP address. A client commonly chooses a temporary ephemeral source port, while a server listens on a known service port.

A transport endpoint can be described by the combination of source IP address, source port, destination IP address, destination port, and transport protocol. For example, a exchange might use the tuple (192.0.2.10, 51514, 198.51.100.20, 443, ). Different client source addresses or source ports allow a server to distinguish multiple clients using the same service port.

When data arrives, uses the protocol and port information to deliver it to the correct socket and application. This is how one host can support web traffic, name-service traffic, and other application traffic simultaneously.

Takeaway: IP addresses locate hosts, ports locate application endpoints, and delivers incoming transport data to the right process.

as a Reliable Byte Stream

provides a connection-oriented, reliable, ordered byte stream. Its reliability comes from several coordinated mechanisms:

  • Sequence numbers identify the positions of bytes in the stream.

  • Acknowledgments report the next sequence number the receiver expects.

  • Retransmissions recover data that appears to have been lost.

  • Duplicate suppression prevents repeated segments from producing repeated application data.

  • Checksums help detect corruption.

  • protects the receiver's buffers.

  • responds to the capacity of the network path.

does not preserve application message boundaries. If an application writes three blocks, the receiver might read them as one block, several blocks, or differently sized pieces. Applications that need records or messages must define boundaries themselves, for example with a length field or delimiter.

Suppose a sender transmits three consecutive byte ranges and the middle range is lost. The receiver can acknowledge the next missing position, while later data may wait until the gap is filled. detects the loss through a retransmission timer, duplicate acknowledgments, or other loss-detection feedback, then retransmits the missing range. The application ultimately receives one ordered stream rather than separate network segments.

Takeaway: makes independent packets appear to be a reliable ordered stream, but applications must add their own message framing when boundaries matter.

Connection Management

A normally establishes a connection. The client sends a SYN with an initial sequence number; the server responds with a SYN-ACK containing its own sequence number and an acknowledgment; the client completes the exchange with an ACK. This synchronizes sequence-number state and confirms that both endpoints can send and receive.

A connection is logical rather than a dedicated physical circuit. Routers still forward individual IP packets independently, while the endpoints maintain the state needed to reconstruct a reliable stream.

shutdown is usually half-closeable. One endpoint can send a FIN when it has no more data in one direction, and the peer can acknowledge that FIN while continuing to send data in the opposite direction. A RST terminates a connection abruptly, such as when traffic refers to a nonexistent connection or an application rejects it. Temporary states after closure help prevent delayed packets from being confused with a later connection.

Takeaway: connection management establishes shared state, supports independent shutdown of each direction, and provides an abrupt reset mechanism when graceful closure is not appropriate.

Flow and

protects the destination host from a sender that transmits faster than the receiver can process or buffer. The receiver advertises a receive window, commonly written as rwndrwnd, indicating how much additional data it can accept. The sender keeps its unacknowledged data within that advertised limit.

If the receive buffer fills, the receiver can advertise a zero window. The sender pauses new transmission and periodically checks whether the receiver has available space again.

addresses a different limitation: the capacity of the network path between the endpoints. maintains a congestion window, commonly written as cwndcwnd. In a simplified model, the usable sending window is

min⁡(rwnd,cwnd)\min(rwnd, cwnd)

Slow start increases the congestion window rapidly at the beginning of a transfer or after a major reduction. Congestion avoidance grows more cautiously as the estimated path capacity is approached. Loss generally causes the sender to reduce its rate, and repeated retransmission timeouts cause increasingly cautious timing. Fast retransmit and fast recovery can respond to multiple duplicate acknowledgments before a timer expires.

These mechanisms serve different purposes. Retransmission repairs detected loss; protects the receiver; reduces pressure on the network and helps avoid congestion collapse.

Takeaway: 's sending rate is constrained by both endpoint capacity and network capacity, represented by the receive window and congestion window.

Datagram Semantics

provides a minimal datagram service. Each send operation produces a separate datagram, and the receiver obtains separate datagram boundaries when delivery succeeds. uses source and destination ports, a length field, and a checksum, but it does not require a transport-level connection handshake.

does not inherently guarantee delivery, retransmit lost datagrams, preserve order, suppress duplicates, provide receiver , or provide general-purpose . A datagram may be lost, duplicated, delayed, or delivered out of order. Applications that need those services must implement them or use a higher-level protocol.

For example, if an application sends datagrams of 100, 200, and 300 bytes, the receiver gets three distinct datagrams if they arrive successfully. It does not get one continuous 600-byte stream. This message-oriented behavior is useful when each update is independently meaningful, but payload sizes must account for path MTU limitations because an oversized datagram may be fragmented, discarded, or fail to reach the application.

checksums provide error detection and help verify delivery to the intended IP addresses and ports. Applications should enable checksums; over IPv6 normally requires one, with only narrowly defined tunnel cases allowing a zero checksum.

A locally connected socket may provide convenient filtering or error reporting, but it does not establish a remote -like connection or require the peer to maintain equivalent connection state.

Takeaway: preserves message boundaries and minimizes transport machinery, but applications must deliberately handle loss, ordering, rate control, and other missing services.

Choosing Between and

and provide contrasting communication abstractions. The most important differences are:

  • Basic abstraction: provides an ordered byte stream; provides independent datagrams.

  • Connection setup: requires connection establishment; has no transport-level handshake.

  • Reliability: includes acknowledgments and retransmission; does not include them by default.

  • Ordering: presents data to the application in order; does not guarantee ordering.

  • Message boundaries: does not preserve application message boundaries; preserves each datagram boundary.

  • : includes receiver ; applications or higher-level protocols must supply it when needed.

  • : includes ; applications or higher-level protocols must supply it when needed.

  • Latency behavior: 's handshake and loss recovery can add delay; has low transport overhead, but the application is responsible for loss recovery.

is a natural fit when every byte must arrive correctly and in order, when the application uses a continuous stream, or when built-in flow and are valuable. Examples include software updates, file transfers, remote terminal sessions, web content, and database results.

is a possible fit when messages are independently meaningful, low latency matters more than perfect delivery, multicast or broadcast is useful, or the application needs control over retransmission, ordering, or timing. Examples include some real-time media, online games, DNS queries, service discovery, and telemetry. These applications still need to control their sending rate and respond appropriately to loss.

illustrates a third design pattern: a higher-level protocol can use as its packet carrier while adding connection state, streams, , loss recovery, , encryption integration, and connection migration. Thus, -based does not necessarily mean unreliable application communication; it means that the additional services are supplied above rather than by .

Takeaway: Choose the protocol according to the required data model, reliability, latency, control, and network-behavior requirements rather than treating as simply slow or as simply fast.