Kubernetes Networking: concept and overview from underlying perspective [programming-RML7]

Kubernetes was built to run distributed systems on a cluster of nodes. Understanding the concept of kubernetes networking could help you correctly understand how to run, monitor and troubleshooting your applications on kubernetes, even more you can know how to choose a suitable distributed system by knowing how to compare them.

To understand it's networking configuration, we have to start from container and how the operating system provides these resource isolations. We start from network namespace this concept of Linux, and create a mock environment for learning how it works as container does. Now, let's begin!

Network namespace [local-3]

Before we start, my environment is Ubuntu 18.04 LTS, and here is the kernel information:

$ uname -a
Linux test-linux 4.15.0-1032-gcp #34-Ubuntu SMP Wed May 8 13:02:46 UTC 2019 x86_64 x86_64 x86_64 GNU/Linux

Create a new network namespace [local-0]

# create network namespace net0
$ ip netns add net0
# create network namespace net1
$ ip netns add net1
# then check
$ ip netns list
net1
net0 (id: 0)

Now we have several network namespaces could emit process on it, but the process can't connect to other networks is meaningless. To solve this problem, we have to create a tunnel for them, in Linux, we can use veth pair to connect two namespaces directly.

Connect two of them with a veth pair [local-1]

# new veth pair
$ ip link add type veth
# assign veth0 to net0
$ ip link set veth0 netns net0
# assign veth1 to net1
$ ip link set veth1 netns net1
$ ip netns exec net0 ip link set veth0 up
# assign ip 10.0.1.2 to veth0, you can use `ip addr` to check it
$ ip netns exec net0 ip addr add 10.0.1.2/24 dev veth0
$ ip netns exec net1 ip link set veth1 up
# assign ip 10.0.1.3 to veth1
$ ip netns exec net1 ip addr add 10.0.1.3/24 dev veth1

NOTE: An important thing is veth pair can't exist alone if you remove one, another would be removed.

Now, ping the network namespace net1 from net0

$ ip netns exec net0 ping 10.0.1.3 -c 3

tcpdump from target network namespace, of course, you should run tcpdump before you ping it.

$ ip netns exec net1 tcpdump -v -n -i veth1
tcpdump: listening on veth1, link-type EN10MB (Ethernet), capture size 262144 bytes
13:54:11.800223 IP6 (hlim 255, next-header ICMPv6 (58) payload length: 16) fe80::905d:ccff:fe4a:cd81 > ff02::2: [icmp6 sum ok] ICMP6, router solicitation, length 16
          source link-address option (1), length 8 (1): 92:5d:cc:4a:cd:81
13:54:12.400440 IP (tos 0x0, ttl 64, id 45855, offset 0, flags [DF], proto ICMP (1), length 84)
    10.0.1.2 > 10.0.1.3: ICMP echo request, id 1433, seq 1, length 64
13:54:12.400464 IP (tos 0x0, ttl 64, id 41348, offset 0, flags [none], proto ICMP (1), length 84)
    10.0.1.3 > 10.0.1.2: ICMP echo reply, id 1433, seq 1, length 64
13:54:13.464163 IP (tos 0x0, ttl 64, id 45912, offset 0, flags [DF], proto ICMP (1), length 84)
    10.0.1.2 > 10.0.1.3: ICMP echo request, id 1433, seq 2, length 64
13:54:13.464189 IP (tos 0x0, ttl 64, id 41712, offset 0, flags [none], proto ICMP (1), length 84)
    10.0.1.3 > 10.0.1.2: ICMP echo reply, id 1433, seq 2, length 64
13:54:14.488184 IP (tos 0x0, ttl 64, id 46671, offset 0, flags [DF], proto ICMP (1), length 84)
    10.0.1.2 > 10.0.1.3: ICMP echo request, id 1433, seq 3, length 64
13:54:14.488221 IP (tos 0x0, ttl 64, id 41738, offset 0, flags [none], proto ICMP (1), length 84)
    10.0.1.3 > 10.0.1.2: ICMP echo reply, id 1433, seq 3, length 64

HTTP can work also. HTTP server:

$ ip netns exec net1 python3 -m http.server
Serving HTTP on 0.0.0.0 port 8000 ...
# After you execute the following command here would show
10.0.1.2 - - [15/May/2019 13:55:41] "GET / HTTP/1.1" 200 -

HTTP client:

$ ip netns exec net0 curl 10.0.1.3:8000
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN" "http://www.w3.org/TR/html4/strict.dtd">
<html>
<head>
<title>Directory listing for /</title>
</head>
...

Although veth pair could help you connect two network namespaces, however, it can't work with more. While we are working on an environment with more than two network namespaces, we would need a more powerful technology: Bridge.

Connect more of them with a bridge [local-2]

# create bridge
$ ip link add br0 type bridge
$ ip link set dev br0 up
# create veth pair for net0, veth0 & veth1
$ ip link add type veth
# create veth pair for net1, veth2 & veth3
$ ip link add type veth
# set up veth pair of net0
$ ip link set dev veth0 netns net0
# You would find veth0 disappeared now by `ip link`
$ ip netns exec net0 ip link set dev veth0 name eth0
$ ip netns exec net0 ip addr add 10.0.1.2/24 dev eth0
$ ip netns exec net0 ip link set dev eth0 up
# bind veth pair of net0 to br0
$ ip link set dev veth1 master br0
$ ip link set dev veth1 up
# set up veth pair of net1
$ ip link set dev veth2 netns net1
$ ip netns exec net1 ip link set dev veth2 name eth0
$ ip netns exec net1 ip addr add 10.0.1.3/24 dev eth0
$ ip netns exec net1 ip link set dev eth0 up
# bind veth pair of net1 to br0
$ ip link set dev veth3 master br0
$ ip link set dev veth3 up

Now, ping 10.0.1.3 from net0 to check our bridge network.

$ ip netns exec net0 ping 10.0.1.3 -c 3
PING 10.0.1.3 (10.0.1.3) 56(84) bytes of data.
64 bytes from 10.0.1.3: icmp_seq=1 ttl=64 time=0.030 ms
64 bytes from 10.0.1.3: icmp_seq=2 ttl=64 time=0.059 ms
64 bytes from 10.0.1.3: icmp_seq=3 ttl=64 time=0.051 ms

--- 10.0.1.3 ping statistics ---
3 packets transmitted, 3 received, 0% packet loss, time 2038ms
rtt min/avg/max/mdev = 0.030/0.046/0.059/0.014 ms

tcpdump our bridge: br0

$ tcpdump -v -n -i br0
tcpdump: listening on br0, link-type EN10MB (Ethernet), capture size 262144 bytes
12:43:39.619458 IP (tos 0x0, ttl 64, id 63269, offset 0, flags [DF], proto ICMP (1), length 84)
    10.0.1.2 > 10.0.1.3: ICMP echo request, id 3046, seq 1, length 64
12:43:39.619553 IP (tos 0x0, ttl 64, id 54235, offset 0, flags [none], proto ICMP (1), length 84)
    10.0.1.3 > 10.0.1.2: ICMP echo reply, id 3046, seq 1, length 64
12:43:40.635730 IP (tos 0x0, ttl 64, id 63459, offset 0, flags [DF], proto ICMP (1), length 84)
    10.0.1.2 > 10.0.1.3: ICMP echo request, id 3046, seq 2, length 64
12:43:40.635764 IP (tos 0x0, ttl 64, id 54318, offset 0, flags [none], proto ICMP (1), length 84)
    10.0.1.3 > 10.0.1.2: ICMP echo reply, id 3046, seq 2, length 64
12:43:41.659714 IP (tos 0x0, ttl 64, id 63548, offset 0, flags [DF], proto ICMP (1), length 84)
    10.0.1.2 > 10.0.1.3: ICMP echo request, id 3046, seq 3, length 64
12:43:41.659742 IP (tos 0x0, ttl 64, id 54462, offset 0, flags [none], proto ICMP (1), length 84)
    10.0.1.3 > 10.0.1.2: ICMP echo reply, id 3046, seq 3, length 64
12:43:44.859619 ARP, Ethernet (len 6), IPv4 (len 4), Request who-has 10.0.1.2 tell 10.0.1.3, length 28
12:43:44.859638 ARP, Ethernet (len 6), IPv4 (len 4), Request who-has 10.0.1.3 tell 10.0.1.2, length 28
12:43:44.859686 ARP, Ethernet (len 6), IPv4 (len 4), Reply 10.0.1.2 is-at 0a:e0:a1:07:b7:c9, length 28
12:43:44.859689 ARP, Ethernet (len 6), IPv4 (len 4), Reply 10.0.1.3 is-at d2:b6:de:2f:4e:f6, length 28

As you thought, br0 would get the traffic from net0 to net1, now we have topology looks like:

bridge_mode_and_namespace.svg

At the final of the output of tcpdump we can see some ARP request/reply, we would talk about it in the next section.

To get more info:

ARP [local-4]

ARP(Address Resolution Protocol) is a communication protocol used for discovering the link layer address, such as a MAC address associated with a given internet layer address.

NOTE: In IPv6(Internet Protocol Version 6), the functionality of ARP provided by NDP(Neighbor Discovery Protocol).

We aren't going to show the whole packet layout of ARP, but mention the part we care in the case.

The working process is:

  1. send ARP request packet with source MAC and source IP and target IP to broadcast address
  2. the machine thought it has this target IP would send ARP reply packet contains it's MAC address
  3. the machine sends ARP request would cache the mapping of IP and MAC into ARP cache, so next time it doesn't have to send ARP request again.

NOTE: others endpoint would ignore non-interested ARP request

arp_request_and_reply.svg

At the previous section, we can see both sides send ARP request to get another IP's information.

To get more info:

Pod to Pod [local-7]

Pod is the unit of kubernetes, so the most basic networking is how to connect from PodA to PodB. We would get two situation:

  1. two Pod at the same Node
  2. two Pod at the different Node

Node is a VM or a machine owned by the kubernetes cluster.

The following discussion assumes a bridge-style setup: every Node runs a Linux bridge, and every Pod gets a veth pair with the host end attached to that bridge. That is the simplest arrangement and it is what most CNI plugins do, so the packet flow below carries over even when the plugin differs.

Pods on the same Node [local-5]

In this case it is just the first section again: the Pods hang off the same bridge, conventionally named cbr0, each through its own veth pair.

Since we already covered this in the first section, we do not spend more time here. The interesting part is how to let Pods on different Nodes reach each other, which is the next section.

pod_to_pod_at_the_same_node.svg

Pods on different Nodes [local-6]

When containers have to talk across hosts the problem shows up. In the traditional model, any container that wanted to reach the outside world had to borrow a port on its host. In a cluster that does not scale, because ports are a limited resource. That is why we have CNM and CNI. We are not going to discuss their details, only the packet flow they have to produce.

Two Pods on different Nodes cannot share a bridge, so the packet cannot simply be forwarded at layer 2.

concept_of_nodes.svg

The whole packet flow looks like:

  1. PodA sends an ARP request
  2. ARP fails, so bridge cbr0 sends the packet out the default route, the host's eth0
  3. routing sends the packet to the default gateway
  4. the default gateway sends it to the right host by CIDR, e.g. 10.0.2.101 to 10.10.0.3
  5. the Node owning PodB's CIDR decides the packet belongs to its own cbr0
  6. cbr0 finally hands the packet to PodB
pod_to_pod_at_different_node_via_default_gateway.svg

Pod to Service [local-11]

We show how to route traffic between Pods and their IP addresses. The model works good until we have to scale the Pod. To make Kubernetes be a great system, we need to have the ability to add/delete resource automatically, which is the main feature of Kubernetes, now problem comes, because we could remove the Pod, we couldn't trust it's IP, since the new Pod won't get the same IP mostly.

To solve the problem, Kubernetes provide an abstraction called Service. A Service had some selectors and some port mappings with a cluster IP, which means it would select Pods as it's backend by selector and to loadbalancing for them and forward packets by port mappings. So whatever how Pods been created or deleted, Service would find those Pods with labels matched selectors, and we only have to know the IP of Service than know all IPs of Pods.

Now, let's take a look at how it works.

iptables and netfilter [local-8]

Kubernetes relies on netfilter – the networking framework bulit-in to Linux.

To get more info about netfilter please take a look at:

iptables is one of userspace tools based on the netfilter providing a table-based system for defining rules for manipulating and transforming packets. In Kubernetes, kube-proxy controller would config iptables rules by watching the changes from API server. The rule monitoring the traffic to the cluster IP of Service and picking a IP from IPs of Pods then forwarding the traffic to the picked IP by updating the destination IP from the cluster IP to the picked IP. This rule would be updated by cluster IP changed, Pod ADDED, Pod DELETED. Which means loadbalancing already been done on the machine to take traffic directed to cluster IP to an actual IP of Pod.

pod_to_service.svg

After the destination IP be updated, the networking model would fall back to the Pod to Pod model.

You can get more details of iptables via:

Load balancing [local-9]

Now I would create some iptables rules to mock a Service for static IPs. Assuming we have three IPs are: 10.244.1.2, 10.244.1.3, 10.244.1.4 and a cluster IP: 10.0.0.2

Start with a single backend:

iptables -t nat -A PREROUTING \
    -p tcp -d 10.0.0.2 --dport 80 \
    -j DNAT --to-destination 10.244.1.2:8080

Reading it piece by piece: -t nat picks the nat table, -A PREROUTING appends to the PREROUTING chain, -p tcp -d 10.0.0.2 --dport 80 matches only TCP traffic aimed at the cluster IP on port 80, and -j DNAT --to-destination 10.244.1.2:8080 rewrites the destination to a Pod.

Unfortunately, we can't just apply this command on to each IPs we want to loadbalance, because the first rule would take all the jobs from others(but our work won't be, damn). That's why iptables provides a module called statistic can work with two different modes:

  • random: probability
  • nth: round robin algorithm

Note: loadbalancing only works during the connection phase of the TCP protocol. Once the connection has been established, the connection would be routed to the same server.

We only introduce round robin here, since it's quite easy to understand and we want to talk about loadbalancing than how loadbalancing works.

$ export CLUSTER_IP=10.0.0.2
$ export SERVICE_PORT=80
$ iptables \
  -A PREROUTING \
  -p tcp \
  -t nat -d $CLUSTER_IP \
  --dport $SERVICE_PORT \
  -m statistic --mode nth \
  --every 3 --packet 0 \
  -j DNAT \
  --to-destination 10.244.1.2:8080

$ iptables \
  -A PREROUTING \
  -p tcp \
  -t nat -d $CLUSTER_IP \
  --dport $SERVICE_PORT \
  -m statistic --mode nth \
  --every 2 --packet 0 \
  -j DNAT \
  --to-destination 10.244.1.3:8080

$ iptables \
  -A PREROUTING \
  -p tcp \
  -t nat -d $CLUSTER_IP \
  --dport $SERVICE_PORT \
  -j DNAT \
  --to-destination 10.244.1.4:8080

The return path [local-10]

The DNAT above only rewrites the destination. The source stays 10.244.1.2, so podB replies to 10.244.1.2 --- an address on its own subnet, which means the reply goes straight back across the bridge at layer 2 instead of being routed. Whether conntrack ever sees that reply decides whether the connection works at all, and that is controlled by one sysctl:

net.bridge.bridge-nf-call-iptables

With it set to 0, bridged frames skip netfilter completely. Watching podA's interface while it talks to the cluster IP:

10.244.1.2.47394 > 10.0.0.2.80:      Flags [S]
10.244.1.3.8080  > 10.244.1.2.47394: Flags [S.]
10.244.1.2.47394 > 10.244.1.3.8080:  Flags [R]

The SYN went to the cluster IP but the SYN-ACK came back from the Pod IP. podA has no socket matching that address pair, so it resets the connection and the request never completes.

The tempting fix is to rewrite the source on the way out:

iptables -t nat -A POSTROUTING \
    -p tcp -s 10.244.1.2 --sport 8080 \
    -j SNAT --to-source 10.0.0.2:80

It does not help. The reply is bridged, so POSTROUTING never sees it either --- with this rule in place the capture is unchanged, still SYN-ACK from the Pod IP followed by a reset. The rule is fixing the wrong layer.

Set the sysctl to 1 and bridged IPv4 traffic is handed to netfilter, which is all conntrack needs:

10.244.1.2.49076 > 10.0.0.2.80:      Flags [S]
10.0.0.2.80      > 10.244.1.2.49076: Flags [S.]
10.244.1.2.49076 > 10.0.0.2.80:      Flags [P.]  HTTP: GET / HTTP/1.1

The reply now arrives from the cluster IP, because NAT in netfilter is connection tracked: the translation is computed once, on the first packet, and conntrack applies the reverse of it to everything coming back. The conntrack entry records both directions:

tcp TIME_WAIT src=10.244.1.2 dst=10.0.0.2 sport=49076 dport=80
              src=10.244.1.3 dst=10.244.1.2 sport=8080 dport=49076 [ASSURED]

The second tuple is the reply it expects, and matching against it is what turns 10.244.1.3:8080 back into 10.0.0.2:80. No SNAT required. This is why a Kubernetes node is expected to have net.bridge.bridge-nf-call-iptables set to 1 --- without it, every Service whose backend Pod sits on the same bridge as its client is quietly broken.

To get more info about loadbalancing and NAT (network address translation):

Internet to Service [local-14]

Egress [local-12]

Egress is traffic from a Pod out to the internet, say a packet from a Pod to some external service: 10.244.1.10 -> 8.8.8.8.

This is not straightforward, because 8.8.8.8 has no idea who 10.244.1.10 is --- they are not on the same network. So we need a globally routable IP, and the rewriting is called masquerading. Assume we have 219.140.7.218: the goal is to turn 10.244.1.10 into 219.140.7.218 before the packet reaches 8.8.8.8, and to turn it back on the way home.

Which raises the next problem: if several Pods make outgoing requests at once, which one gets a given reply? The simple answer (there are other NAT schemes) is to allocate a port per connection. 10.244.1.10:8080 -> 8.8.8.8:53 is rewritten as 219.140.7.218:61234, and a concurrent 10.244.1.11:8080 -> 8.8.8.8:53 as 219.140.7.218:61235. The replies come back to different ports, so each can be rewritten back to the right Pod.

Ingress [local-13]

A load balancer is the easy case: it just provides an IP for your Service and does exactly what the internal cluster IP does --- rewrite the destination and send the packet to the right Pod.

An ingress controller is an application layer load balancer instead. It terminates the HTTP request itself and routes on the path:

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: hello-world-ingress
  annotations:
    nginx.ingress.kubernetes.io/ssl-redirect: "false"
    nginx.ingress.kubernetes.io/rewrite-target: /$1
spec:
  ingressClassName: nginx
  rules:
    - http:
        paths:
          - path: /hello(/|$)(.*)
            pathType: ImplementationSpecific
            backend:
              service:
                name: hello-svc
                port:
                  number: 80
          - path: /world(/|$)(.*)
            pathType: ImplementationSpecific
            backend:
              service:
                name: world-svc
                port:
                  number: 80

The controller owns the root path / and dispatches by prefix: anything under /hello goes to hello-svc, anything under /world goes to world-svc. Note that the controller usually picks the backend Pod inside its own code rather than going through the Service's cluster IP, so the Service here mostly serves as the list of endpoints.