diff --git a/docs/user-guide/host-settings.md b/docs/user-guide/host-settings.md index b381e8df..d45a5c62 100644 --- a/docs/user-guide/host-settings.md +++ b/docs/user-guide/host-settings.md @@ -107,42 +107,82 @@ As a first step, users should download the docker image from our registry: docker pull ghcr.io/githedgehog/host-bgp ``` +The command above downloads the `latest` image for the host-bgp container; alternatively, a specific +version can be fetched by adding it as a suffix to the image name, e.g., using version `v0.4.0` as a reference: + +```bash +docker pull ghcr.io/githedgehog/host-bgp:v0.4.0 +``` + +Pinning a version is recommended, so that a host reboot cannot silently pick up a different image; the examples +below all refer to `v0.4.0`. + +!!! note + Starting with `v0.4.0`, the image is published as a multi-arch manifest covering `linux/amd64` and `linux/arm64`, + so the same image name works on both x86 and ARM hosts and `docker pull` picks the matching variant automatically. + Earlier versions are `amd64`-only. + The container should then be run with host networking (so that FRR can communicate with the leaves using the host's interfaces) and in privileged mode. Additionally, a few input parameters are required: - an optional ASN to use in BGP. If present, it should be the first parameter; if not specified, the container will use ASN 64999 - one or more VPC subnets with their related parameters, in the format - `:v=:i=[:i=...]:a=[:a=...]`, where: + `:v=:i=[,][:i=[,]...]:a=[:a=...]`, where: - `` is just a mnemonic ID for the VPC subnet we want to attach to. It can be anything as long as it is a legal name for a route-map or prefix-list in FRR. - `v=` is the VLAN ID to be used for the VPC; use 0 for untagged. - - `i=` is an interface to be used to establish a BGP unnumbered session with a + - `i=` is an interface to be used to establish a BGP unnumbered session with a Fabric leaf; if a VLAN ID was specified, a corresponding VLAN interface will be created using - the provided interface as the master device. - - `a=` is the Virtual IP (or VIP) to be advertised to the leaves; it should have + the provided interface as the master device, named `.` by default. + - `` is an optional name for that VLAN interface, given after a comma separator, and used + instead of the default `.`. It must be a legal interface name (alphanumeric + characters, dashes and underscores, at most 15 characters) and it is only accepted when the VLAN + ID is non-zero. An alias is *required* when the default `.` name would exceed + the kernel's 15-character limit on interface names: the container refuses to start rather than + truncating the name. + - `a=
` is the Virtual IP (or VIP) to be advertised to the leaves; it should have a prefix length of /32 and be part of the subnet the host is attaching to. - an optional list of static routes (which, if present, must go after the VPC subnets), in the format `--routes :i=[:d=] [:i=[:d=]...]`, where: - `` is the destination IP prefix for the route. - `i=` is the egress interface for the route. - `d=` is an optional administrative distance for the route, in the range 1-255 (default 1). +- an optional list of per-interface rail IPs, useful to run RoCE from the host (which, if present, must also go after + the VPC subnets), in the format + `--roce :a= [:a=...]`, where: + - `` refers to an interface declared in the VPC section. In the `--roce` section, reference it + either by its real interface name or by its alias (if one was configured), but not using the comma form + (`,`). The interface must appear in exactly one VPC subnet, since the rail IP rides + that subnet's existing BGP session, and it can carry a single rail IP. + - `a=
` is a /32 IP address that will be configured on the enslaved VLAN interface (or the "parent" + interface itself if not using VLANs) and advertised via the BGP session over that interface only, so that + traffic towards it follows a single link. Note that by default Fabric leaves will only accept /32 prefixes + belonging to the hostBGP subnet proper; if these per-interface addresses are not part of that, appropriate + prefixes including them should be configured as `hostBGPExtraPrefixes` in the + [hostBGP subnet](vpcs.md#hostbgp-subnets). + +The `--routes` and `--roce` sections can appear in any order relative to each other, as long as both come after the +VPC subnets. As an example, the command might look something like this: ```bash -docker run --network=host --privileged --rm --detach --name hostbgp ghcr.io/githedgehog/host-bgp 64307 vpc-01:v=1001:i=enp2s1:i=enp2s2:a=10.100.34.5/32 --routes 0.0.0.0/0:i=enp2s0:d=100 +docker run --network=host --privileged --rm --detach --name hostbgp ghcr.io/githedgehog/host-bgp:v0.4.0 64307 vpc-01:v=1001:i=enp2s1:i=enp2s2,to-leaf2:a=10.100.34.5/32 --routes 0.0.0.0/0:i=enp2s0:d=100 --roce enp2s1:a=10.115.1.5/32 to-leaf2:a=10.115.2.7/32 ``` !!! note With the above command, any output produced by the container will not be visible from the terminal - where it was started. Verify that the container is running correctly with `docker ps`, or examine - the logs of the container with `docker logs hostbgp` to investigate a failure. + where it was started. Verify that the container is running correctly with `docker ps` and inspect its output + with `docker logs hostbgp`. Note that if the container fails at startup, `--rm` will have already removed it and + its logs along with it; in that case, drop both `--detach` and `--rm` to see the failure directly. With the above command: -- VLAN interfaces `enp2s1.1001` and `enp2s2.1001` would be created, if not already existing +- VLAN interfaces `enp2s1.1001` and `to-leaf2` would be created, if not already existing - BGP unnumbered sessions would be created on those same interfaces, using ASN 64307 - the address `10.100.34.5/32` would be configured on the loopback of the host server and it would be advertised to the leaves - a default route via `enp2s0` would be installed with administrative distance 100; this is higher than the default distance for a BGP route (20), meaning that any default route learned via BGP would take priority over this, effectively making it a "backup" route +- IP address `10.115.1.5/32` would be configured on `enp2s1.1001` and advertised over the corresponding BGP session +- IP address `10.115.2.7/32` would be configured on `to-leaf2` and advertised over the corresponding BGP session To further modify the configuration or to troubleshoot the state of the system, an expert user can invoke the FRR CLI using the following command: @@ -163,28 +203,56 @@ hostname server-04 service integrated-vtysh-config ! ip prefix-list vpc-01 seq 5 permit 10.100.34.5/32 +ip prefix-list roce-enp2s1-v1001 seq 5 permit 10.115.1.5/32 +ip prefix-list roce-to-leaf2-v1001 seq 5 permit 10.115.2.7/32 ! route-map vpc-01 permit 10 match ip address prefix-list vpc-01 exit ! +route-map roce-enp2s1-v1001 permit 10 + match ip address prefix-list vpc-01 +exit +! +route-map roce-enp2s1-v1001 permit 20 + match ip address prefix-list roce-enp2s1-v1001 +exit +! +route-map roce-to-leaf2-v1001 permit 10 + match ip address prefix-list vpc-01 +exit +! +route-map roce-to-leaf2-v1001 permit 20 + match ip address prefix-list roce-to-leaf2-v1001 +exit +! ip route 0.0.0.0/0 enp2s0 100 ! +interface enp2s1.1001 + ip address 10.115.1.5/32 +exit +! interface lo ip address 10.100.34.5/32 exit ! +interface to-leaf2 + ip address 10.115.2.7/32 +exit +! router bgp 64307 no bgp ebgp-requires-policy bgp bestpath as-path multipath-relax timers bgp 3 9 neighbor enp2s1.1001 interface remote-as external - neighbor enp2s2.1001 interface remote-as external + neighbor to-leaf2 interface remote-as external ! address-family ipv4 unicast network 10.100.34.5/32 - neighbor enp2s1.1001 route-map vpc-01 out - neighbor enp2s2.1001 route-map vpc-01 out + network 10.115.1.5/32 + network 10.115.2.7/32 + neighbor enp2s1.1001 route-map roce-enp2s1-v1001 out + neighbor to-leaf2 route-map roce-to-leaf2-v1001 out maximum-paths 4 exit-address-family exit @@ -204,7 +272,7 @@ be removed manually; for example, using iproute2 and the reference command above sudo ip route delete 0.0.0.0/0 dev enp2s0 sudo ip address delete dev lo 10.100.34.5/32 sudo ip link delete dev enp2s1.1001 -sudo ip link delete dev enp2s2.1001 +sudo ip link delete dev to-leaf2 ``` Users should consider automating the startup of the hostbgp container at system boot up, to make @@ -249,7 +317,7 @@ core@control-1 ~ $ kubectl fabric vpc attach --name=s3-v2-l2 --conn=server-03--u Then we can configure `server-03` using the provided container: ```bash -docker run --network=host --privileged --rm --detach --name hostbgp ghcr.io/githedgehog/host-bgp vpc-01:v=1001:i=enp2s1:i=enp2s2:a=10.0.1.3/32 vpc-02:v=1002:i=enp2s1:i=enp2s2:a=10.0.2.3/32 +docker run --network=host --privileged --rm --detach --name hostbgp ghcr.io/githedgehog/host-bgp:v0.4.0 vpc-01:v=1001:i=enp2s1:i=enp2s2:a=10.0.1.3/32 vpc-02:v=1002:i=enp2s1:i=enp2s2:a=10.0.2.3/32 ``` This will generate the following FRR configuration: diff --git a/docs/user-guide/vpcs.md b/docs/user-guide/vpcs.md index ad6decd6..3111f3b3 100644 --- a/docs/user-guide/vpcs.md +++ b/docs/user-guide/vpcs.md @@ -60,6 +60,15 @@ spec: bgp-on-host: # Another subnet with hosts peering with leaves via BGP subnet: 10.10.50.0/25 hostBGP: true + hostBGPMinPrefixLen: 25 # optional shortest prefix length accepted from the host within the subnet's IP range, default is 32; cannot be shorter than the subnet's own prefix length + hostBGPMaxPrefixLen: 32 # optional longest prefix length accepted from the host within the subnet's IP range, default (and maximum) is 32 + hostBGPExtraPrefixes: # optional map of extra IPv4 prefixes, on top of the subnet's own IP range, to accept from the host's BGP sessions (up to 100) + "10.10.55.0/24": # keys must be canonical IPv4 prefixes and must not overlap the fabric's reserved subnets + minPrefixLen: 26 # optional shortest prefix length accepted within the parent prefix, defaults to hostBGPMinPrefixLen + maxPrefixLen: 32 # optional longest prefix length accepted within the parent prefix, defaults to hostBGPMaxPrefixLen + "10.10.56.14/32": # a single host route; minPrefixLen must be set here, since the inherited 25 would be shorter than the prefix itself + minPrefixLen: 32 + maxPrefixLen: 32 vlan: 1050 permit: # Defines which subnets of the current VPC can communicate to each other, applied on top of subnets "isolated" flag (doesn't affect VPC peering) @@ -203,11 +212,22 @@ for example because of hardware limitations. Consider this scenario: `server-1` is connected to two different Fabric switches `sw-1` and `sw-2`, and attached to `vpc-1/subnet-1` on both of them. This subnet is configured as `hostBGP`; the switches will be configured to peer with -`server-1` using unnumbered BGP (IPv4 unicast address family), only importing /32 prefixes in the subnet of the VPC and -exporting routes learned from other VPC peers. Similarly, BGP is running on `server-1`, unnumbered BGP sessions are +`server-1` using unnumbered BGP (IPv4 unicast address family), by default only importing /32 prefixes in the subnet of the +VPC and exporting routes learned from other VPC peers. Similarly, BGP is running on `server-1`, unnumbered BGP sessions are established with each leaf, and one or more Virtual IPs (VIPs) in the VPC subnet are advertised. With this setup, the host is part of the VPC and can be reached via one of the advertised VIPs from either link to the Fabric. +The set of prefixes the leaves accept from the host can be widened in two ways. `hostBGPMinPrefixLen` and +`hostBGPMaxPrefixLen` relax the prefix lengths accepted *within* the subnet's own IP range, so the host can advertise +aggregates rather than only /32s. `hostBGPExtraPrefixes` accepts prefixes *outside* the subnet's IP range altogether, +each with its own optional prefix length bounds. + +The latter is what makes it possible to run RoCE from the host: each interface connecting the host to a leaf is given +its own rail IP, advertised only over that interface's BGP session so that traffic to it follows a single link rather +than being load-balanced. Those rail IPs typically live outside the VPC subnet, so the prefixes covering them must be +listed under `hostBGPExtraPrefixes` for the leaves to accept them. See +[the `--roce` parameter of the HostBGP container](host-settings.md#hostbgp-container) for the host-side configuration. + It is important to keep in mind that Hedgehog Fabric does not directly operate the host servers attached to it; running subnets in HostBGP mode requires running a routing suite and configuring it accordingly. To facilitate this process, however, we do provide a container image which can autogenerate a valid configuration, given some input parameters.