add
- [<ifaceGroup>,]{cf|smpr|ecds|...},<ifaceList> | Creates a new interface group or adds interfaces to an
+ detection.Table 2. Interface (and Interface Group) Commands add
+
[<ifaceGroup>,]
+
{cf|smpr|ecds|...},
+
<ifaceList> | Creates a new interface group or adds interfaces to an
existing one and configures the interface group for a specific
relay algorithm (e.g., cf,
smpr, ecds). Also, a
@@ -373,8 +380,9 @@
flooding or Elastic Multicast routing domain. (Note that
nrlsmf is currently limited to a single
"elastic" interface group but will allow for multiple, separate
- "elastic" groups in the future. | remove
- <ifaceGroup>,[<iface1>[,<iface2>,...]] | Removes either an entire
+ "elastic" groups in the future. | remove
+
<ifaceGroup>,
+
[<iface1>[,<iface2>,...]] | Removes either an entire
<ifaceGroup> or the given
<ifaceList> interfaces from a specific
<group> or from all groups if the
@@ -423,8 +431,9 @@
nrlsmf during run-time. (default =
"off"). In the future, this command will be updated to apply to
specific interface group(s). At the moment, it applies to all
- interface groups. | allow [vrf,<srcVRF>,<dstVRF>,]{all |
- <addr1>[,<addr2>,...]} | Sets a forwarding policy to allow either specific Virtual
+ interface groups. | allow
+
[vrf,<srcVRF>,<dstVRF>,]
+
{all | <addr1>[,<addr2>,...]} | Sets a forwarding policy to allow either specific Virtual
Routing and Forwarding (VRF) groups with specific destination
addresses or globally allow specific addresses to be
flooded/routed by nrlsmf. Non-matching
@@ -435,8 +444,9 @@
provides a way to edit the overall policy to exclude specific
addresses. The allow and
deny commands can be used in combination to
- control the overall policy. | deny [vrf,<srcVRF>,<dstVRF>,]{all |
- <addr1>[,<addr2>,...]} | Sets a forwarding policy to deny (block) destinations from
+ control the overall policy. | deny
+
[vrf,<srcVRF>,<dstVRF>,]
+
{all | <addr1>[,<addr2>,...]} | Sets a forwarding policy to deny (block) destinations from
forwarding within either specific Virtual Routing and Forwarding
(VRF) groups or globally. Matching addresses will be ignored by
forwarding within the VRF group (if 'vrf' is
@@ -445,9 +455,9 @@
. The complementary allow command provides a
way to edit the overall policy to exclude specific addresses. The
allow and deny commands can
- be used in combination to control the overall policy. | device
- <vifName>,<ifaceName>[,<addrList
- ...>] | This command instantiations a tun/tap virtual interface
+ be used in combination to control the overall policy. | device
+
<vifName>,<ifaceName>
+
[,<addrList ...>] | This command instantiations a tun/tap virtual interface
named <vifName> that is associated
with the underlay physical interface identified by
<interfaceName>. If the optional
@@ -457,9 +467,11 @@
reassigned to the new <vifName>
interface. Those addresses and routes are restored to the
<ifaceName> interface when
- nrlsmf shuts down. | cid <vifName>,<iface1>[/{t|r|d}][,
- <iface2>[/{t|r|d}][,<iface3>[/{t|r|d}],...]]
- | The Composite Interface Device (cid)
+ nrlsmf shuts down. | cid
+
<vifName>,
+
<iface1>[/{t|r|d}][,
+
<iface2>[/{t|r|d}][,
+
<iface3>[/{t|r|d}],...]] | The Composite Interface Device (cid)
command can be used to add (or reconfigure) physical interface
devices to an existing 'parent' virtual interface created by the
nrlsmf device
@@ -487,8 +499,8 @@
interface identified by <ifaceName>.
The default behavior is unlimited rate. Note a
<bits/sec> value of -1.0 can be used
- to indicate unlimited rate. | queue
- [<ifaceName>,]<queueLimit> | This command enables an nrlsmf
+ to indicate unlimited rate. | queue
+
[<ifaceName>,]<queueLimit> | This command enables an nrlsmf
layer of queuing in addition to whatever queuing (e.g., Linux
'tc' traffic control) is enabled on the
underlying physical interface. Note that
@@ -512,8 +524,13 @@
designated IP multicast <groupAddr>
on the given <ifaceName>. This is
generally needed for GRE decapsulation to successfully work with
- mGRE operation. This command supports administrative management of
- those necessary memberships. Future
+ multicast-underlay mGRE (the kernel tunnel remote is the group).
+ On a wildcard-remote mGRE device the kernel does not demux
+ GRE-in-multicast by outer destination; if that same group is
+ also a mapped inject remote, ujoin additionally
+ enables an underlay capture so
+ nrlsmf can strip GRE and treat the
+ inner packet as inbound on the overlay GRE interface. Future
nrlsmf Elastic Multcast operation may
support dynamic underlay group membership depending on the overlay
multicast -> underlay multicast mapping scheme used. Meanwhile
@@ -533,7 +550,7 @@
duplicate packet detection, flow detection and timeout) to streamline the
code and its performance. Some aspects of Elastic Multicast design are
still in progress informed by limited deployment and experimentation
- activities.Table 3. Elastic Multicast Commands elastic <ifaceGroup> | Enables Elastic Multicast routing instead of basic SMF
+ activities.Table 3. Elastic Multicast Commands elastic <ifaceGroup> | Enables Elastic Multicast routing instead of basic SMF
flooding for the given <ifaceGroup>.
IMPORTANT - note that nrlsmf is currently
limited to providing proper Elastic Multicast routing on a single
@@ -574,9 +591,10 @@
Thus non-ETC interfaces are assumed lossless with a one-hop cost
of one and mix of ETX and non-ETX interfaces can be used. Elastic
Multicast will be extended with additional metric options in the
- future. | flow
- <srcAddr>[/<maskLen>]->]<dstAddr>[/<maskLen>]
- [,<protocol>[,<class>] | Administratively emplace a flow description that will be
+ future. | flow
+
<srcAddr>[/<maskLen>]->]
+
<dstAddr>[/<maskLen>]
+
[,<protocol>[,<class>] | Administratively emplace a flow description that will be
proactively advertised when Elastic Multicast advertise mode is
enabled. Note a local interface name can be used in place of the
<srcAddr> parameter and the IP
@@ -596,9 +614,10 @@
wildcard addressing. An IP <protocol>
value of 255 will wildcard that aspect of the flow description and
a zero-value <class> does the same
- for its aspect. | join
- <srcAddr>[/<maskLen>]->]<dstAddr>[/<maskLen>]
- [,<protocol>[,<class>] | Administratively "join" a flow description that will prompt
+ for its aspect. | join
+
<srcAddr>[/<maskLen>]->]
+
<dstAddr>[/<maskLen>]
+
[,<protocol>[,<class>] | Administratively "join" a flow description that will prompt
Elastic Multicast routing to acknowledge (via EM_ACK) matching
upstream flows. Nominally, this can simply be an IP multicast
<dstAddr> group or a more specific
@@ -606,24 +625,54 @@
option) flow classifier, even including unicast if
nrlsmf unicast
operation is enabled. The same field specifications apply as for
- the flow command. | leave
- <srcAddr>[/<maskLen>]->]<dstAddr>[/<maskLen>]
- [,<protocol>[,<class>] | Administratively "leave" a flow previously set with the
+ the flow command. | leave
+
<srcAddr>[/<maskLen>]->]
+
<dstAddr>[/<maskLen>]
+
[,<protocol>[,<class>] | Administratively "leave" a flow previously set with the
join command. This will cancel acknowledgment
- of matching flows for Elastic Multicast operation. | map
- <iface>,<localAddr>[,<remoteAddr>] | Maps tunnel endpoint address information for a GRE
+ of matching flows for Elastic Multicast operation. | map
+
<iface>,<localAddr>
+
[,<remoteAddr>] | Maps tunnel endpoint address information for a GRE
interface (or other ancillary iface->address associations).
This enables Elastic Multicast operation to properly match inbound
EM_ACK messages to the interface upon which the corresponding
- EM_ADV or forwarded packet was transmitted. Typically, this
- command is not needed when nrlsmf is able to
- automatically retrieve the tunnel endpoint information when a
- tunnel interface is added to an interface group. However, some
- tunnel configurations (e.g. mGRE) may require this "helper"
- function as part of nrlsmf setup. | unmap
- <iface>,<localAddr>[,<remoteAddr>] | Deletes matching tunnel endpoint or interface->address
+ EM_ADV or forwarded packet was transmitted. Repeat the command
+ with the same local address and a different remote to record
+ additional remotes; overlay multicast forwarded onto a
+ wildcard-remote mGRE interface is transmitted once to each
+ mapped remote that is a unicast peer or an underlay multicast
+ group. A remote of dynamic learns unicast
+ remotes from the kernel neighbor table on <iface> and
+ keeps them updated (NHRP or static NBMA). It does not apply to
+ metadata/"external" GRE, which needs explicit unicast remotes
+ for overlay-multicast inject. A remote of
+ 0.0.0.0 is the kernel wildcard
+ (INADDR_ANY), not a send destination. Typically this command
+ is not needed for point-to-point GRE or multicast-underlay
+ mGRE (the tunnel remote is already the group). Runtime mappings
+ can be listed with "show tunnel":
+ Local/Remote are underlay
+ tunnel endpoints and IP is the local overlay
+ address. C in
+ Flags means the mapping was added with this
+ command; otherwise it was learned from the kernel. GRE/mGRE
+ neighbors are listed with "show tunnel
+ neighbors": Neighbor IP is the peer
+ overlay address and Remote is the peer
+ underlay. Neighbor flags use C for
+ configured plus the kernel NUD state:
+ R reachable, S stale,
+ D delay, P probe,
+ I incomplete, F failed,
+ N no-ARP, M
+ permanent (for example CR is configured and
+ reachable). | unmap
+
<iface>,<localAddr>
+
[,<remoteAddr>] | Deletes matching tunnel endpoint or interface->address
associations previously set up with the map
- command. | reliable <ifaceList> | Enable experimental reliable hop-by-hop forwarding on
+ command. A remote of dynamic stops neighbor
+ learning and removes learned remotes; explicit unicast
+ map entries are left in place. | reliable <ifaceList> | Enable experimental reliable hop-by-hop forwarding on
listed interfaces (adds UMP option to IPv4 packets). | utos <trafficClass> | Set the IP traffic class value to be ignored by reliable
forwarding. | adaptive <group> | Enable Smart Routing (adaptive routing) for the named
interface group. |
The "Gateway Commands" apply to configuring
@@ -638,15 +687,15 @@
The nrlsmf application has some logic in its
command-parsing to avoid some of these scenarios, but the flexibility
allowed permits some configurations that may produce undesirable
- effects. Table 4. Gateway Commands push
+ effects.Table 4. Gateway Commands push
<srcIface>,<dstIfaceList> | Force relay of packets from
<srcIface> to listed destination
interfaces. The <dstIfaceList> is a
comma-delimited (no spaces) list of interface names. Packet
forwarding is subject to duplicate packet detection and TTL/hop
limit constraints. Useful for setting up an SMF "gateway" to
- inject packets to/from a MANET SMF area. | rpush
- <srcIface>,<dstIfaceList> | Resequence (modify IPv4 ID field or add IPv6 SMF-DPD option
+ inject packets to/from a MANET SMF area. | rpush
+
<srcIface>,<dstIfaceList> | Resequence (modify IPv4 ID field or add IPv6 SMF-DPD option
header) and force relay of packets from
<srcIface> to listed destination
interfaces. The <dstIfaceList> is a
@@ -678,64 +727,481 @@
packet identifier "resequencing" (or "tagging") performed here is
for "forwarded" (i.e. relayed) packets and is distinct from the
"resequence" command that applies to
- locally-generated IP Multicast packets. |
nrlsmf supports operation on Linux Generic
- Routing Encapsulation (GRE) interfaces, including point-to-point GRE and
- point-to-multipoint (mGRE) tunnels. GRE interfaces are treated
- differently from Ethernet interfaces: packets received on a GRE interface
- may not include a usable Ethernet header, and tunnel endpoint addresses
- (local and remote) are used for forwarding and elastic multicast control
- traffic. When a GRE interface is opened, nrlsmf
- attempts to read tunnel local and remote endpoint addresses from the
- operating system. If valid endpoint information is available, it is
- recorded automatically. Some GRE deployments (for example, lightweight
- or externally-managed "metadata" GRE tunnels) may not expose complete
- endpoint information to user space. In those cases, the
- "map" command must be used to associate tunnel
- endpoint addresses with the interface. A warning is logged when a GRE
- interface is used without adequate endpoint mapping. For mGRE operation, the remote endpoint may correspond to
- INADDR_ANY (0.0.0.0). The underlay network may
- require joining a multicast group so encapsulated traffic can be
- received; use the "ujoin" and
- "uleave" commands for this purpose on the physical
- underlay interface (not on the GRE interface itself). When elastic multicast is enabled, GRE tunnel endpoints are also
- used to identify the correct interface for elastic control messages
- (for example, EM_ACK). JSON configuration file support for GRE tunnel
- mapping is not yet available; use command-line or run-time remote
- control commands instead. Table 5. GRE Tunnel Commands map
- <iface>,<localAddr>[,<remoteAddr>] | Associate tunnel endpoint address information with
- <iface>. If
- <remoteAddr> is omitted or invalid, the
- <localAddr> is recorded as an ordinary
- local interface address. If
- <remoteAddr> is valid, the mapping is
- recorded as tunnel endpoint information for GRE or other tunnel
- interfaces. For mGRE, <remoteAddr> may
- be 0.0.0.0. | unmap
- <iface>,<localAddr>[,<remoteAddr>] | Remove a previously configured tunnel or local address
- mapping from <iface>. The address
- arguments must match those supplied to
- "map". | ujoin
- <groupAddr>,<iface> | Join multicast group <groupAddr>
- on underlay interface <iface> to enable
- reception of mGRE multicast-encapsulated traffic. The group
- address must be a valid IP multicast address. | uleave
- <groupAddr>,<iface> | Leave multicast group <groupAddr>
- on underlay interface <iface>. |
Example: The following example adds GRE and
- Ethernet interfaces to a flooding group, maps explicit tunnel endpoints
- for a metadata GRE interface, and joins an underlay multicast group
- for mGRE reception: nrlsmf add mygroup,cf,gre0,eth0 \
- map gre0,10.0.0.1,10.0.0.2 \
- ujoin 239.1.1.1,eth0 Virtual Interface CommandsSMF supports virtual interface constructs that associate a
+ locally-generated IP Multicast packets. |
nrlsmf supports operation on Linux
+ Generic Routing Encapsulation (GRE) and point-to-multipoint (mGRE)
+ interfaces. From a configuration standpoint, a GRE or mGRE interface
+ is just another interface name that can be placed into an
+ nrlsmf interface group alongside physical
+ interfaces. Overlay unicast peer resolution (kernel
+ ip neigh, NHRP, or a multicast underlay group)
+ stays outside nrlsmf. Overlay
+ multicast inject onto a wildcard-remote mGRE
+ interface is the exception:
+ nrlsmf transmits once per mapped remote
+ (unicast peers and, optionally, an underlay multicast group).
+ Overlay GRE interfaces are typically
+ layered so a packet received on the tunnel is
+ not flooded back out of it. The NHRP hub is the exception: it must
+ replicate overlay multicast to the spokes. In every mode
+ nrlsmf still floods decapsulated overlay
+ multicast per its relay rules and keeps endpoint information for
+ duplicate packet detection (DPD) and Elastic Multicast. GRE interfaces are treated differently from Ethernet interfaces:
+ packets received on a GRE interface may not include a usable Ethernet
+ header, and tunnel endpoint addresses (local and remote) are used for
+ forwarding and Elastic Multicast control traffic. Endpoint Address DiscoveryWhen a GRE interface is opened,
+ nrlsmf reads the tunnel's local and
+ remote endpoint addresses from the operating system. This is enough
+ for point-to-point GRE and for multicast-underlay mGRE (the
+ tunnel remote is already a multicast group): there is a single
+ remote, so one transmission has a well-defined destination. map
+ <iface>,<localAddr>[,<remoteAddr>]
+ records tunnel endpoint information in
+ nrlsmf's interface-info table. Multiple
+ map commands with the same local address and
+ different remotes add multiple peers. That list is used in two
+ ways:
DPD / Elastic Multicast bookkeeping (all GRE modes). Overlay multicast inject onto a wildcard-remote
+ (0.0.0.0) mGRE interface: the kernel has
+ no single destination for overlay multicast, so
+ nrlsmf transmits once per mapped
+ remote that is a unicast peer or an underlay multicast group.
+ Unicast remotes come from explicit
+ map <iface>,<local>,<peer>
+ and/or
+ map <iface>,<local>,dynamic
+ (kernel neighbor table). A mapped remote may also be an
+ underlay multicast group; that is an explicit
+ map, not learned by
+ dynamic, and inbound GRE-in-multicast on
+ that group needs ujoin as well. Overlay
+ unicast on that same
+ interface still uses the kernel neighbor table
+ (ip neigh), not map.
+ map gre0,10.0.0.2,0.0.0.0 is allowed: it
+ records the kernel wildcard remote
+ (INADDR_ANY) for DPD / Elastic Multicast
+ bookkeeping. It is not a send destination and not a synonym
+ for dynamic. Opening a tunnel that already
+ has remote 0.0.0.0 stores the same
+ mapping, so that command is usually redundant.
The other case that requires map is a
+ "lightweight" or "metadata" GRE interface (externally-managed
+ tunnel with no fixed endpoints on the device).
+ nrlsmf logs a warning that the
+ interface is missing tunnel endpoint addressing until
+ map supplies it. When Elastic Multicast is enabled, GRE tunnel endpoints are
+ also used to identify the correct interface for elastic control
+ messages (for example, EM_ACK). JSON configuration file support
+ for GRE tunnel mapping is not yet available; use command-line or
+ run-time remote control commands instead. A point-to-point GRE tunnel has one fixed local address and
+ one fixed remote address, configured once at tunnel setup. There
+ is no ambiguity about where an outbound packet goes, so no
+ ongoing peer-resolution mechanism is needed. Topology: Node A WAN / Underlay Node B
+ +---------------+ +---------------+
+ | eth0 (uplink)|<------------------------------------------->| eth0 (uplink)|
+ | 10.0.0.2 | | 10.0.1.2 |
+ | gre0 |=========== GRE tunnel (P2P) ================| gre0 |
+ | 172.16.0.1 | | 172.16.0.2 |
+ | wlan0 (MANET)| | eth1 (LAN) |
+ | SMF flooding | | multicast |
+ | | | receivers |
+ +---------------+ +---------------+ Kernel-level setup (Node A): ip tunnel add gre0 mode gre local 10.0.0.2 remote 10.0.1.2 ttl 32
+ip addr add 172.16.0.1/30 dev gre0
+ip link set gre0 up nrlsmf configuration (Node A): nrlsmf resequence on \
+ cf wlan0 \
+ rpush wlan0,gre0 \
+ push gre0,wlan0 \
+ layered gre0 cf wlan0 — classical flooding within
+ the MANET.
rpush wlan0,gre0 — relay
+ MANET-originated traffic across the tunnel, tagging it for
+ DPD.
push gre0,wlan0 — relay traffic from
+ the far side back into the MANET (already tagged).
layered gre0 — a point-to-point
+ tunnel gains nothing from being re-flooded out itself.
Typical use case: connecting two MANET
+ "islands," or a MANET to a fixed-infrastructure multicast network,
+ across a WAN hop with no native multicast support. A single mGRE interface represents multiple remote peers
+ instead of one. GRE itself has no built-in signaling for which
+ peer a packet goes to; three different mechanisms answer that
+ question for overlay unicast, entirely below
+ nrlsmf. Modes A and B are, at the kernel level, the same mechanism
+ — an mGRE interface with no fixed remote
+ (remote 0.0.0.0) and a neighbor-resolution
+ table that maps overlay tunnel addresses to underlay peer
+ addresses — differing only in whether that table is maintained by
+ a routing daemon or by hand. Mode C is fundamentally different:
+ instead of per-peer unicast resolution, it turns the tunnel into a
+ virtual broadcast/multicast segment. Mode A — NHRP-resolved mGRE (dynamic hub-and-spoke)Peers ("spokes") register their real underlay address with
+ a Next Hop Resolution Protocol (NHRP) server (typically a hub,
+ or "NHS" — Next Hop Server). Any peer needing to reach another
+ resolves the destination's tunnel-address-to-underlay-address
+ mapping dynamically through the NHS, and the kernel encapsulates
+ directly to the resolved address. This is the mechanism behind
+ Cisco DMVPN-style deployments; on Linux/FRR it is provided by
+ FRR's nhrpd daemon (RFC 2332), which
+ handles NHRP registration and resolution independently of
+ nrlsmf.
+ nrlsmf itself has no NHRP support
+ and does not need any: by the time a decapsulated packet
+ reaches nrlsmf on the mGRE
+ interface, nhrpd and the kernel have
+ already resolved and delivered it. Topology: Hub / NHRP Server (NHS)
+ +------------------------+
+ | eth0 (uplink) |
+ | 10.0.0.2 |
+ | gre0 (mGRE) |
+ | 172.16.0.1 |
+ +-----------+------------+
+ |
+ Underlay / WAN
+ dynamic NHRP resolution, per-peer unicast
+ / | \
+ +-------------------+ +-------------------+ +-------------------+
+ | Spoke S1 | | Spoke S2 | | Spoke S3 |
+ | eth0 10.0.1.2 | | eth0 10.0.2.2 | | eth0 10.0.3.2 |
+ | gre0 172.16.0.2 | | gre0 172.16.0.3 | | gre0 172.16.0.4 |
+ +-------------------+ +-------------------+ +-------------------+ Kernel + FRR setup (Hub): ip tunnel add gre0 mode gre local 10.0.0.2 key 42 ttl 64
+ip addr add 172.16.0.1/32 dev gre0
+ip link set gre0 up Omitting remote is the same as
+ remote 0.0.0.0 (kernel wildcard): ip tunnel add gre0 mode gre local 10.0.0.2 remote 0.0.0.0 key 42 ttl 64 interface gre0
+ ip nhrp network-id 1
+ ip nhrp registration no-unique The hub's own ordinary underlay address
+ (10.0.0.2, whatever interface it already
+ uses to reach the WAN) doubles as its NBMA identity — no
+ separate loopback or special addressing is required. Kernel + FRR setup (Spoke S1; S2 and S3 use
+ 10.0.2.2/172.16.0.3 and
+ 10.0.3.2/172.16.0.4): ip tunnel add gre0 mode gre local 10.0.1.2 key 42 ttl 64
+ip addr add 172.16.0.2/32 dev gre0
+ip link set gre0 up Equivalent wildcard remote: ip tunnel add gre0 mode gre local 10.0.1.2 remote 0.0.0.0 key 42 ttl 64 interface gre0
+ ip nhrp network-id 1
+ ip nhrp nhs 172.16.0.1 nbma 10.0.0.2
+ ip nhrp registration no-unique nhrpd installs overlay-to-
+ underlay neighbors. A spoke's neighbor table is the NHS; the
+ hub's table is every registered spoke. Overlay unicast therefore
+ looks like static NBMA. Overlay multicast does not: a spoke has
+ only the hub as a mapped unicast remote, so overlay multicast
+ from a spoke goes to the hub, and the hub's
+ nrlsmf must replicate it to the
+ other spokes. Spokes are layered so a packet
+ received on the tunnel is not flooded back out of it. The hub
+ is not layered, because it is the overlay
+ replicator. nrlsmf configuration (hub; each spoke
+ underlay address is a remote): nrlsmf add overlay,cf,eth1,gre0 \
+ map gre0,10.0.0.2,10.0.1.2 \
+ map gre0,10.0.0.2,10.0.2.2 \
+ map gre0,10.0.0.2,10.0.3.2 Equivalent using neighbor learning
+ (nhrpd installs the same unicast
+ remotes in ip neigh): nrlsmf add overlay,cf,eth1,gre0 \
+ map gre0,10.0.0.2,dynamic nrlsmf configuration (spoke S1; S2 and S3 use
+ their own underlay address): nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.1.2,dynamic map …,dynamic is the usual inject
+ config with NHRP:
+ nhrpd programs
+ ip neigh, and
+ nrlsmf learns those unicast remotes.
+ Explicit maps of 0.0.0.0 record the kernel
+ wildcard remote for DPD / Elastic Multicast bookkeeping; they
+ are not send destinations and not a synonym for
+ dynamic.
Production DMVPN deployments typically add IPsec
+ (nhrpd integrates with strongSwan
+ via the VICI protocol) to protect the overlay traffic; that is
+ independent of nrlsmf and omitted
+ here. Spoke-to-spoke "shortcut" tunnels (DMVPN Phase 2/3)
+ require additional NFLOG/iptables configuration and are not
+ used for overlay multicast with
+ nrlsmf classic flooding — see FRR's
+ nhrpd documentation for that
+ setup. Mode B — Statically-mapped mGRE (manual NBMA table)Mechanically identical to Mode A at the kernel level —
+ same remote 0.0.0.0 mGRE interface — but
+ the peer-resolution table is populated by hand instead of by a
+ routing daemon. Overlay unicast is resolved by the kernel neighbor table
+ (overlay-addr to underlay-addr). On Node A: ip neigh add 172.16.0.2 lladdr 10.0.1.2 dev gre0 nud permanent
+ip neigh add 172.16.0.3 lladdr 10.0.2.2 dev gre0 nud permanent
+ip neigh add 172.16.0.4 lladdr 10.0.3.2 dev gre0 nud permanent Overlay multicast has no entry in that table (and a
+ single send cannot fan out to several unicast underlay
+ destinations). nrlsmf injects
+ overlay multicast once per mapped remote in its
+ interface-info table. Those remotes are usually the same
+ unicast peers as in ip neigh, but a mapped
+ remote may also be an underlay multicast group (see
+ the section called “Mixed underlay: multicast group plus unicast peers”). The topology has the same shape as Mode A's diagram, with
+ static ip neigh entries in place of
+ nhrpd; typically used for a small,
+ fixed set of known peers. nrlsmf configuration (Node A, local
+ underlay 10.0.0.2): nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.0.2,10.0.1.2 \
+ map gre0,10.0.0.2,10.0.2.2 \
+ map gre0,10.0.0.2,10.0.3.2 Equivalent using neighbor learning from the
+ ip neigh table above
+ (dynamic dumps that table and keeps it
+ updated via NEWNEIGH/DELNEIGH): nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.0.2,dynamic The wildcard remote can be mapped explicitly. This is
+ allowed, but it is not a send destination and not a synonym
+ for dynamic. Opening
+ gre0 already records it from the kernel,
+ so it is usually redundant: nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.0.2,0.0.0.0 Typical use case: a handful of
+ fixed, stable sites where the operational overhead of running
+ NHRP is not worth it, but a single mGRE interface (instead of
+ one point-to-point interface per peer) is still
+ convenient. Mode C — Multicast-underlay mGREThe tunnel's remote address is configured as an IP
+ multicast group address rather than a unicast peer or a
+ wildcard. A single encapsulated transmission is sent once, to
+ that group; the underlay network's own multicast routing (for
+ example, PIM) replicates and delivers it to every peer that
+ has joined the group. This is the mode
+ nrlsmf's
+ ujoin/uleave commands
+ exist for, and it is the most natural fit for SMF: symmetric,
+ no hub, no per-peer replication — the same job SMF already
+ does on a physical broadcast medium, relocated onto a
+ WAN. Topology: Underlay network with native IP multicast
+ routing (e.g. PIM), group 239.1.1.1
+ / | \
+ / | \
+ +-----------------+ +-----------------+ +-----------------+
+ | Node A | | Node B | | Node C |
+ | eth0 (underlay) | | eth0 (underlay) | | eth0 (underlay) |
+ | 10.0.0.2 | | 10.0.1.2 | | 10.0.2.2 |
+ | gre0 (mGRE) | | gre0 (mGRE) | | gre0 (mGRE) |
+ | 172.16.0.1 | | 172.16.0.2 | | 172.16.0.3 |
+ | remote=239.1.1.1| | remote=239.1.1.1| | remote=239.1.1.1|
+ | wlan0 (MANET) | | wlan0 (MANET) | | wlan0 (MANET) |
+ +-----------------+ +-----------------+ +-----------------+ Kernel-level setup (Node A; Node B and Node C
+ use 10.0.1.2/172.16.0.2
+ and
+ 10.0.2.2/172.16.0.3 —
+ symmetric, no hub role): ip tunnel add gre0 mode gre local 10.0.0.2 remote 239.1.1.1 ttl 16
+ip addr add 172.16.0.1/24 dev gre0
+ip link set gre0 up nrlsmf configuration (each node): nrlsmf add overlay,cf,gre0,wlan0 \
+ layered gre0 \
+ ujoin 239.1.1.1,eth0 ujoin/uleave are
+ issued against the underlay interface
+ (eth0), not the GRE interface itself — they
+ join or leave the multicast group on the physical network so
+ the kernel GRE device (whose remote is already that group)
+ actually receives the encapsulated traffic to decapsulate.
+ Mirror with uleave 239.1.1.1,eth0 on
+ shutdown or reconfiguration. This is not the same as mapping
+ a multicast group onto a wildcard-remote mGRE device; the
+ kernel already demuxes GRE-in-multicast onto the tunnel, so
+ nrlsmf does not also strip GRE from
+ underlay capture (that would double-process the packet). When Elastic Multicast routing is enabled, the same
+ tunnel endpoint information is used to route EM control
+ traffic (for example, EM_ACK) to the
+ correct interface — another reason accurate endpoint state
+ matters even when data-plane decapsulation "just works." This mode only works if the underlay genuinely supports
+ IP multicast routing end-to-end. If it does not, GRE-over-
+ multicast can silently fall back to resolving individual
+ neighbors instead of true multicast fan-out — worth verifying
+ the underlay's multicast path explicitly before relying on
+ this mode. Mixed underlay: multicast group plus unicast peersWildcard-remote mGRE (Modes A and B) can inject overlay
+ multicast to an underlay multicast group as well as to unicast
+ peers. That is useful when some sites can join the group and
+ others cannot: one transmission covers the multicast-capable
+ set, and additional transmissions cover the unicast-only
+ peers. The kernel tunnel remote stays
+ 0.0.0.0; this is not multicast-underlay
+ mGRE (Mode C), whose device remote is already the group. A Linux wildcard-remote mGRE device cannot send overlay
+ multicast as GRE-in-multicast by itself (there is no multicast
+ sll_addr for
+ PF_PACKET), and it does not demux inbound
+ GRE-in-multicast either: kernel GRE lookup matches the tunnel
+ local address, not the outer multicast destination. Mapping
+ the group as an inject remote makes
+ nrlsmf send GRE-in-multicast once
+ to that group. ujoin of the same group on
+ the underlay interface both joins IGMP and enables an underlay
+ capture so nrlsmf can strip GRE and
+ treat the inner packet as inbound on the overlay GRE
+ interface. Do not enable that capture on a multicast-underlay
+ GRE device: the kernel already delivers the packet on the
+ tunnel, and a second strip would double-process it. Topology: five overlay routers on
+ one wildcard-remote mGRE. Nodes A, B, and C can join underlay
+ group 239.1.1.1. Nodes D and E cannot. Node A -- lan0 --\
+ Node B -- lan1 ---\
+ Node C -- lan2 ---- underlay A/B/C join 239.1.1.1
+ Node D -- lan3 ---/ D/E unicast only
+ Node E -- lan4 --/ Kernel-level setup (Node A; B–E use their own
+ underlay and overlay addresses). Same wildcard-remote mGRE as
+ Modes A and B — the tunnel remote is
+ 0.0.0.0, not the underlay group: ip tunnel add gre0 mode gre local 10.0.0.2 remote 0.0.0.0 ttl 64
+ip addr add 172.16.0.1/24 dev gre0
+ip link set gre0 multicast on
+ip link set gre0 up Overlay unicast still needs kernel neighbors (Node A;
+ these can also be added by a dynamic protocol such as
+ NHRP): ip neigh add 172.16.0.2 lladdr 10.0.1.2 dev gre0 nud permanent
+ip neigh add 172.16.0.3 lladdr 10.0.2.2 dev gre0 nud permanent
+ip neigh add 172.16.0.4 lladdr 10.0.3.2 dev gre0 nud permanent
+ip neigh add 172.16.0.5 lladdr 10.0.4.2 dev gre0 nud permanent On A, B, and C, point the underlay group at the physical
+ iface so GRE-in-multicast has an output device: ip route replace 239.1.1.1/32 dev eth0 nrlsmf configuration (Node A; B and C are the
+ same with their own underlay address): nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ ujoin 239.1.1.1,eth0 \
+ map gre0,10.0.0.2,239.1.1.1 \
+ map gre0,10.0.0.2,10.0.3.2 \
+ map gre0,10.0.0.2,10.0.4.2 Overlay multicast inject from A is then three GRE
+ transmissions: one to 239.1.1.1, one to D
+ (10.0.3.2), and one to E
+ (10.0.4.2). Overlay unicast still uses
+ kernel ip neigh. Nodes D and E create the same wildcard-remote tunnel
+ (Node D: local 10.0.3.2 / overlay
+ 172.16.0.4) and the same style of
+ ip neigh entries. They do not add the
+ underlay group route. nrlsmf configuration (Node D; E is the same
+ with its own underlay address and no
+ ujoin): nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.3.2,10.0.0.2 \
+ map gre0,10.0.3.2,10.0.1.2 \
+ map gre0,10.0.3.2,10.0.2.2 \
+ map gre0,10.0.3.2,10.0.4.2 layered on the overlay GRE interface
+ keeps a packet that arrived on the tunnel from being flooded
+ back out of it, so a unicast-only router does not re-inject
+ toward the multicast set. Traffic that arrives on
+ eth1 is still injected onto
+ gre0. A fourth way GRE encapsulation parameters can be resolved,
+ distinct from all three mGRE modes above:
+ ip link add ... type gre external creates an
+ interface with no fixed local, remote, or key at all. Instead,
+ whatever adds routes or flow rules for it supplies the
+ encapsulation parameters per destination, using Linux's
+ lightweight tunnel ("lwtunnel") route encap: ip route add 172.16.0.2/32 encap ip id 42 src 10.0.0.2 dst 10.0.1.2 ttl 16 dev gre0 This is the mechanism SDN controllers and OVS/OVN typically
+ use to build many per-flow or per-peer tunnels dynamically out of
+ a single device. It is conceptually the multipoint-resolution
+ equivalent of Mode B's static neighbor table — same job,
+ expressed as routes instead of neighbor entries — but because
+ the device itself has no fixed endpoint addresses, this is the
+ one case where nrlsmf's
+ map command is actually mandatory: there is
+ nothing for nrlsmf to auto-discover
+ from the interface, and it will log a warning
+ (GRE interface '<name>' missing tunnel endpoint
+ addressing. (must map it)) until map
+ supplies a local address. Overlay unicast still uses the kernel
+ lwtunnel routes. Overlay multicast inject has no single remote,
+ so nrlsmf also needs unicast remotes
+ in that same map table
+ (map <iface>,<local>,<peer>
+ per peer).
+ map …,0.0.0.0 is only the wildcard
+ bookkeeping form; map …,dynamic does not
+ apply (no ip neigh). nrlsmf configuration (Node A): nrlsmf add overlay,cf,gre0 \
+ layered gre0 \
+ map gre0,10.0.0.2,10.0.1.2 \
+ map gre0,10.0.0.2,10.0.2.2 Table 5. GRE / mGRE modes | Mode | Peer resolution | nrlsmf extras | Typical use |
|---|
| P2P GRE | Fixed at setup | typically layered | Two fixed sites; MANET-island bridging over a WAN
+ hop | | mGRE Mode A (NHRP) | Dynamic, per-peer unicast (FRR
+ nhrpd) | map …,dynamic (or explicit
+ unicast maps); spokes
+ layered, hub not (replicator) | Hub-and-spoke overlay multicast; NHRP/DMVPN-style
+ tooling already in place | | mGRE Mode B (static NBMA) | Static, per-peer unicast (manual
+ ip neigh) | explicit unicast maps (or
+ map …,dynamic); typically
+ layered | Small, fixed set of known peers; avoids running
+ NHRP | | mGRE mixed underlay | Same wildcard-remote device as Modes A/B;
+ overlay multicast inject to an underlay group
+ and unicast peers | ujoin plus
+ map of the group (GRE-in-mcast send
+ and underlay demux); unicast
+ maps for peers that cannot join;
+ typically layered | Some sites can join underlay multicast, others
+ cannot | | mGRE Mode C (multicast underlay) | None — underlay multicast fan-out (PIM) | ujoin/uleave
+ on the underlay interface (IGMP join; kernel demuxes
+ GRE-in-mcast onto the tunnel) | Symmetric SMF/MANET-gateway overlays | | External (metadata) GRE | Per-destination lwtunnel routes (often
+ SDN-managed) | map is mandatory; remotes for
+ overlay multicast inject; typically
+ layered | SDN/OVS-managed overlays; an external controller
+ provisions tunnel endpoints |
Modes A and B are the same underlying mechanism — an
+ mGRE interface with no fixed remote and a peer-resolution
+ table — differing only in whether that table is maintained
+ dynamically (NHRP) or by hand. The same wildcard-remote
+ device can also map an underlay multicast group as one
+ overlay-multicast inject dest (mixed underlay); that is not
+ multicast-underlay mGRE (Mode C). Point-to-point and multicast-underlay mGRE (Mode C)
+ rely on nrlsmf's automatic
+ endpoint discovery (one configured remote). Modes A and B
+ overlay unicast use kernel neigh
+ (nhrpd or
+ ip neigh); overlay multicast inject uses
+ SMF's map table, filled by explicit
+ maps (unicast peers and/or an underlay
+ multicast group) and/or
+ map …,dynamic. map is mandatory for
+ "external"/metadata GRE (no endpoints on the device). For
+ Mode A/B overlay multicast, mapped remotes (or
+ map …,dynamic)
+ are required. 0.0.0.0 means wildcard
+ remote, matching the kernel. It is not needed for
+ point-to-point or Mode C. Mapping an underlay multicast
+ group onto a wildcard-remote mGRE also requires
+ ujoin of that group so inbound
+ GRE-in-multicast can be demuxed. In every mode, once packets reach
+ nrlsmf they are handled by the
+ same relay logic
+ (cf / smpr /
+ ecds / push /
+ rpush / merge /
+ rmerge / layered)
+ as any other interface. Overlay multicast inject onto a
+ wildcard-remote mGRE is the extra send-side step: one
+ transmission per mapped remote (unicast peer or underlay
+ group), still using that same relay decision. Overlay GRE
+ interfaces are typically layered; the
+ NHRP hub is not, because it replicates overlay multicast.
Table 6. GRE Tunnel Commands map
+
<iface>,<localAddr>
+
[,<remoteAddr>] | Associate tunnel endpoint address information with
+ <iface>. If
+ <remoteAddr> is omitted or
+ invalid, the <localAddr> is
+ recorded as an ordinary local interface address. If
+ <remoteAddr> is valid, the
+ mapping is recorded as tunnel endpoint information for
+ GRE or other tunnel interfaces. For mGRE,
+ <remoteAddr> may be 0.0.0.0
+ (kernel wildcard remote, not a send dest).
+ dynamic learns unicast remotes from
+ the kernel neighbor table (NHRP or static NBMA). It does
+ not apply to metadata/"external" GRE. Multiple commands
+ with different remotes record multiple inject
+ destinations (unicast peers and/or an underlay multicast
+ group) on a wildcard-remote mGRE interface or a metadata
+ GRE device. | unmap
+
<iface>,<localAddr>
+
[,<remoteAddr>] | Remove a previously configured tunnel or local
+ address mapping from
+ <iface>. The address arguments
+ must match those supplied to
+ "map". A remote of
+ dynamic stops neighbor learning and
+ removes learned remotes only. | ujoin
+ <groupAddr>,<iface> | Join multicast group
+ <groupAddr> on underlay
+ interface <iface> so the host
+ receives GRE-in-multicast. On multicast-underlay mGRE
+ this is an IGMP join so the kernel tunnel (remote
+ already the group) can decapsulate. On a wildcard-remote
+ mGRE device, if that same group is also a mapped inject
+ remote, this additionally enables underlay capture so
+ nrlsmf can strip GRE and
+ treat the inner packet as inbound on the overlay GRE
+ interface. The group address must be a valid IP
+ multicast address. | uleave
+ <groupAddr>,<iface> | Leave multicast group
+ <groupAddr> on underlay
+ interface <iface>. |
Virtual Interface CommandsSMF supports virtual interface constructs that associate a
user-space forwarding path with one or more underlying physical
interfaces. These are commonly used for host integration, composite
- devices, and encapsulation. Table 6. Virtual Interface Commands device
- <vifName>,<ifaceName>[/{t|r|d}][,<addr>...] | Create a virtual SMF interface (vif) associated with one
+ devices, and encapsulation.Table 7. Virtual Interface Commands device
+
<vifName>,<ifaceName>
+
[/{t|r|d}][,<addr>...] | Create a virtual SMF interface (vif) associated with one
or more physical interfaces. Optional flags on
<ifaceName> are
t (transmit), r (receive),
or d (delete). Additional addresses may be
- listed after the interface name. | cid
- <vifName>,<iface1>[/{t|r|d}][,<iface2>...] | Add or remove elements of a Composite Interface Device
+ listed after the interface name. | cid
+
<vifName>,
+
<iface1>[/{t|r|d}][,
+
<iface2>...] | Add or remove elements of a Composite Interface Device
associated with an SMF virtual interface. | layered <ifaceList> | Mark interfaces as layered, indicating they have their
own underlying multicast distribution mechanism. | igmpProxy <ifaceList> | Send IGMP joins on listed interfaces for groups of
interest. Typically used with layered interfaces. | encapsulate <ifaceList> | Use IPIP encapsulation for outbound unicast packets on
@@ -743,14 +1209,15 @@
device operation. (default = on) | route
<dstAddr>,<nextHopAddr> | Add a debug route used with encapsulation testing. |
nrlsmf can associate interfaces with named
virtual routing and forwarding (VRF) contexts. When run alongside FRR,
- VRF information may be imported at startup. Table 7. VRF Commands vrf
- <vrf-name>,[<vrf-id>,]<ifaceList> | Associate one or more interfaces with the named VRF. An
+ VRF information may be imported at startup.Table 8. VRF Commands vrf
+
<vrf-name>,
+
[<vrf-id>,]<ifaceList> | Associate one or more interfaces with the named VRF. An
optional numeric VRF identifier may be supplied. | with-frr | Run alongside FRR and pull configuration where appropriate,
such as VRF information. This option must appear early on the
command line. |
nrlsmf supports loading and saving
configuration in JSON format. Some advanced options (including GRE tunnel
mapping) are not yet available via JSON and must be configured using
- command-line or remote control commands. Table 8. Configuration Commands load <configFile> | Load an nrlsmf JSON configuration file at startup or at
+ command-line or remote control commands.Table 9. Configuration Commands load <configFile> | Load an nrlsmf JSON configuration file at startup or at
run-time. | save <configFile> | Save the current configuration to a JSON file, typically
upon exit. | stats | Return interface and group information via the remote
control interface. |
The "Remote-only Commands" listed here can be invoked only via the
@@ -766,8 +1233,8 @@
list. The "selectorMac" and "neighborMac" apply to control of S-MPR
forwarding. The "mneBlockMac" command is provided to support proper SMF
behavior in the NRL
- Mobile Network Emulation (MNE) environment. Table 9. Remote-only Commands selectorMac <binary
- macAddrArray> | This command can be used by external processes (e.g.,
+ Mobile Network Emulation (MNE) environment.Table 10. Remote-only Commands selectorMac
+
<binary macAddrArray> | This command can be used by external processes (e.g.,
nrlolsrd) to control
nrlsmf S-MPR forwarding. This command
sets the list of MAC addresses that have selected the local
@@ -777,8 +1244,8 @@
addresses. Thus, the length of this "array" is a multiple of 6
bytes. An ASCII "space" character delimits the literal
"selectorMac" command string from the binary
- array of MAC addresses. | neighborMac <binary
- macAddrArray> | This command can be used by external processes (e.g.,
+ array of MAC addresses. | neighborMac
+
<binary macAddrArray> | This command can be used by external processes (e.g.,
nrlolsrd) to control
nrlsmf S-MPR forwarding. This command
sets the list of MAC addresses that have been identified as
@@ -788,8 +1255,8 @@
6-byte Ether-type MAC addresses. Thus, the length of this
"array" is a multiple of 6 bytes. An ASCII "space" character
delimits the literal "selectorMac" command
- string from the binary array of MAC addresses. | mneBlockMac <binary
- macAddrArray> | This command enables the
+ string from the binary array of MAC addresses. | mneBlockMac
+
<binary macAddrArray> | This command enables the
nrlsmf process to be compatible with
the NRL Mobile Network Emulator (MNE) system that uses MAC-based
blocking to emulate the connectivity of mobile network
@@ -859,7 +1326,87 @@
"\\.\mailslot\<instanceName>" is created and
used while on WinCE systems a semaphore is instantiated along with a
corresponding registry entry mapping to a locally-bound UDP socket
- provides equivalent functionality.There are limitations in the use of some
+ provides equivalent functionality. The same nrlsmf binary can act as a
+ client against that socket. Invoke it as
+ "nrlsmf --cli [-i <instanceName>] -c "show <command>
+ [modifiers]"" to query runtime status. Repeat
+ -c to send more than one command (for example
+ "nrlsmf --cli -c "show statistics" -c "show interface grouping"").
+ Commands include
+ version, statistics,
+ interface, interface grouping
+ (a subcommand of interface),
+ tunnel, tunnel neighbors
+ (GRE/mGRE neighbors from configuration and/or the kernel),
+ groups, groups memberships,
+ and igmp groups. The last three require an Elastic
+ Multicast build. The unmodified command is the default listing.
+ Optional modifiers are command-specific. brief requests
+ less output and details requests more; a command may
+ support none, one, or both. When JSON is wanted, json
+ is always the last modifier (for example
+ "show groups brief json").
+ Older one-shot verbs such as stats and
+ groupsj still work for compatibility. Configuration
+ commands (for example "nrlsmf --cli -c "debug 2"",
+ "with-frr", or "elastic overlay")
+ are sent without waiting for a response. Use "-i" when
+ the target process was started with a non-default
+ instance name. Run "nrlsmf --cli ?"
+ for the show-command list. On an Elastic build, show interface includes
+ flag M (JSON Managed) when
+ last-hop membership skip is in effect on that interface. The other
+ flags are L layered, T tunnel,
+ I IGMP proxy, and S
+ shadowing. show groups lists detected multicast flows.
+ The fwd column is the flow default (often
+ L, LIMIT, about 1 packet per second).
+ Per-interface FORWARD after an EM_ACK is not that column; use
+ show groups details.
+ recv is the lifetime packet count.
+ pps is over a tumbling window of about 5
+ seconds.
nrlsmf --cli -i smf-r1 -c "show groups"
+Flags: A = Active, K = Ack
+Fwd: B = Block, H = Hybrid, L = Limit, F = Forward, D = Deny
+group saddr iif flags fwd recv pps
+---------------- ---------------- ------------ ----- --- ------ ----
+239.0.0.1 192.168.55.2 gre1 AK L 41 8.8 show groups details nests inbound sources
+ and outbound interfaces. Selected
+ Y is the current upstream.
+ Why is managed (IGMP or static
+ last hop), ack (EM_ACK from a downstream relay),
+ or -.
nrlsmf --cli -i smf-r1 -c "show groups details"
+Flags: A = Active, K = Ack
+Selected: Y = current upstream, N = additional inbound
+Fwd: B = Block, H = Hybrid, L = Limit, F = Forward, D = Deny
+Why: managed = IGMP/static last hop, ack = EM_ACK, - = default
+
+group saddr flags ack recv pps
+---------------- ---------------- ----- --- ------ ----
+239.0.0.1 192.168.55.2 AK yes 41 8.8
+ sources
+ iif upstream sel
+ gre1 10.0.0.2 Y
+ downstream
+ oif mode why fwd sent pps relays
+ eth1 elastic managed F 41 8.8 -
+ gre1 elastic ack F 0 0.0 - show groups memberships is the controller
+ table: S static (join),
+ M managed (IGMP via with-frr),
+ E elastic (downstream EM_ACK). This is not the
+ last-hop skip list.
nrlsmf --cli -i smf-r1 -c "show groups memberships"
+Flags: S = Static, M = Managed, E = Elastic
+group saddr iface flags relays
+---------------- ---------------- ------------ ----- ----------------
+239.0.0.1 * eth1 M -
+239.0.0.1 192.168.55.2 gre1 E 10.0.0.2 show igmp groups is the last-hop list
+ filled by with-frr. A managed host interface that
+ does not list the group is not used as a last hop.
nrlsmf --cli -i smf-r1 -c "show igmp groups"
+Flags: M = Managed (last-hop HasActiveMembership)
+iface flags groups
+------------ ----- ------
+eth1 M 239.0.0.1
+gre1 - - There are limitations in the use of some
nrlsmf options. Many of these limitations are a
result of nrlsmf being a cross-platform,
user-space implementation. Many of these subtleties could be overcome with
diff --git a/doc/nrlsmf.pdf b/doc/nrlsmf.pdf
index 61b5f1c..8057695 100644
Binary files a/doc/nrlsmf.pdf and b/doc/nrlsmf.pdf differ
diff --git a/doc/nrlsmf.xml b/doc/nrlsmf.xml
index d57d3ac..db2f32f 100644
--- a/doc/nrlsmf.xml
+++ b/doc/nrlsmf.xml
@@ -167,7 +167,7 @@ make -f Makefile.linux elastic
Run "nrlsmf -help" for the complete list of supported commands.
To send a command to an already-running instance:
-nrlsmf --cli [-i <instanceName>] <command> [args...]
+nrlsmf --cli [-i <instanceName>] -c "<command>" [-c "<command>" ...]
@@ -346,9 +346,9 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]Operating Mode Commands
-
+
-
+
@@ -389,7 +389,8 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- hash {MD5 | CRC32 | SHA1 | NONE}
+ hash
+ {MD5 | CRC32 | SHA1 | NONE}
When I-DPD is disabled ("idpd off") and
a hashing algorithm is enabled via this command (i.e.,
@@ -415,8 +416,8 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- ihash {MD5 | CRC32 | SHA1 |
- NONE}
+ ihash
+ {MD5 | CRC32 | SHA1 | NONE}
This command is similar to the "hash"
command described above in that a hashing
@@ -617,8 +618,8 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- dscpCapture
- <dscpValue>,<dscpValueList>
+ dscpCapture
+ <dscpValue>,<dscpValueList>
Specify a list of DSCP values for unicast packets that will
be intercepted by nrlsmf. This command
@@ -658,8 +659,9 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- vrf
- <vrf-name>,[<vrf-id>,]<ifaceList>
+ vrf
+ <vrf-name>,
+ [<vrf-id>,]<ifaceList>
Associate list of interfaces with a VRF.
@@ -680,14 +682,16 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]Interface (and Interface Group) Commands
-
+
-
+
- add
- [<ifaceGroup>,]{cf|smpr|ecds|...},<ifaceList>
+ add
+ [<ifaceGroup>,]
+ {cf|smpr|ecds|...},
+ <ifaceList>
Creates a new interface group or adds interfaces to an
existing one and configures the interface group for a specific
@@ -710,8 +714,9 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- remove
- <ifaceGroup>,[<iface1>[,<iface2>,...]]
+ remove
+ <ifaceGroup>,
+ [<iface1>[,<iface2>,...]]
Removes either an entire
<ifaceGroup> or the given
@@ -796,8 +801,9 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- allow [vrf,<srcVRF>,<dstVRF>,]{all |
- <addr1>[,<addr2>,...]}
+ allow
+ [vrf,<srcVRF>,<dstVRF>,]
+ {all | <addr1>[,<addr2>,...]}
Sets a forwarding policy to allow either specific Virtual
Routing and Forwarding (VRF) groups with specific destination
@@ -814,8 +820,9 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- deny [vrf,<srcVRF>,<dstVRF>,]{all |
- <addr1>[,<addr2>,...]}
+ deny
+ [vrf,<srcVRF>,<dstVRF>,]
+ {all | <addr1>[,<addr2>,...]}
Sets a forwarding policy to deny (block) destinations from
forwarding within either specific Virtual Routing and Forwarding
@@ -830,9 +837,9 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- device
- <vifName>,<ifaceName>[,<addrList
- ...>]
+ device
+ <vifName>,<ifaceName>
+ [,<addrList ...>]
This command instantiations a tun/tap virtual interface
named <vifName> that is associated
@@ -848,9 +855,11 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- cid <vifName>,<iface1>[/{t|r|d}][,
- <iface2>[/{t|r|d}][,<iface3>[/{t|r|d}],...]]
-
+ cid
+ <vifName>,
+ <iface1>[/{t|r|d}][,
+ <iface2>[/{t|r|d}][,
+ <iface3>[/{t|r|d}],...]]
The Composite Interface Device (cid)
command can be used to add (or reconfigure) physical interface
@@ -890,8 +899,8 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- queue
- [<ifaceName>,]<queueLimit>
+ queue
+ [<ifaceName>,]<queueLimit>
This command enables an nrlsmf
layer of queuing in addition to whatever queuing (e.g., Linux
@@ -923,8 +932,13 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]<groupAddr>
on the given <ifaceName>. This is
generally needed for GRE decapsulation to successfully work with
- mGRE operation. This command supports administrative management of
- those necessary memberships. Future
+ multicast-underlay mGRE (the kernel tunnel remote is the group).
+ On a wildcard-remote mGRE device the kernel does not demux
+ GRE-in-multicast by outer destination; if that same group is
+ also a mapped inject remote, ujoin additionally
+ enables an underlay capture so
+ nrlsmf can strip GRE and treat the
+ inner packet as inbound on the overlay GRE interface. Future
nrlsmf Elastic Multcast operation may
support dynamic underlay group membership depending on the overlay
multicast -> underlay multicast mapping scheme used. Meanwhile
@@ -962,9 +976,9 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]Elastic Multicast Commands
-
+
-
+
@@ -1027,9 +1041,10 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- flow
- <srcAddr>[/<maskLen>]->]<dstAddr>[/<maskLen>]
- [,<protocol>[,<class>]
+ flow
+ <srcAddr>[/<maskLen>]->]
+ <dstAddr>[/<maskLen>]
+ [,<protocol>[,<class>]
Administratively emplace a flow description that will be
proactively advertised when Elastic Multicast advertise mode is
@@ -1055,9 +1070,10 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- join
- <srcAddr>[/<maskLen>]->]<dstAddr>[/<maskLen>]
- [,<protocol>[,<class>]
+ join
+ <srcAddr>[/<maskLen>]->]
+ <dstAddr>[/<maskLen>]
+ [,<protocol>[,<class>]
Administratively "join" a flow description that will prompt
Elastic Multicast routing to acknowledge (via EM_ACK) matching
@@ -1071,9 +1087,10 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- leave
- <srcAddr>[/<maskLen>]->]<dstAddr>[/<maskLen>]
- [,<protocol>[,<class>]
+ leave
+ <srcAddr>[/<maskLen>]->]
+ <dstAddr>[/<maskLen>]
+ [,<protocol>[,<class>]
Administratively "leave" a flow previously set with the
join command. This will cancel acknowledgment
@@ -1081,28 +1098,57 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- map
- <iface>,<localAddr>[,<remoteAddr>]
+ map
+ <iface>,<localAddr>
+ [,<remoteAddr>]
Maps tunnel endpoint address information for a GRE
interface (or other ancillary iface->address associations).
This enables Elastic Multicast operation to properly match inbound
EM_ACK messages to the interface upon which the corresponding
- EM_ADV or forwarded packet was transmitted. Typically, this
- command is not needed when nrlsmf is able to
- automatically retrieve the tunnel endpoint information when a
- tunnel interface is added to an interface group. However, some
- tunnel configurations (e.g. mGRE) may require this "helper"
- function as part of nrlsmf setup.
+ EM_ADV or forwarded packet was transmitted. Repeat the command
+ with the same local address and a different remote to record
+ additional remotes; overlay multicast forwarded onto a
+ wildcard-remote mGRE interface is transmitted once to each
+ mapped remote that is a unicast peer or an underlay multicast
+ group. A remote of dynamic learns unicast
+ remotes from the kernel neighbor table on <iface> and
+ keeps them updated (NHRP or static NBMA). It does not apply to
+ metadata/"external" GRE, which needs explicit unicast remotes
+ for overlay-multicast inject. A remote of
+ 0.0.0.0 is the kernel wildcard
+ (INADDR_ANY), not a send destination. Typically this command
+ is not needed for point-to-point GRE or multicast-underlay
+ mGRE (the tunnel remote is already the group). Runtime mappings
+ can be listed with "show tunnel":
+ Local/Remote are underlay
+ tunnel endpoints and IP is the local overlay
+ address. C in
+ Flags means the mapping was added with this
+ command; otherwise it was learned from the kernel. GRE/mGRE
+ neighbors are listed with "show tunnel
+ neighbors": Neighbor IP is the peer
+ overlay address and Remote is the peer
+ underlay. Neighbor flags use C for
+ configured plus the kernel NUD state:
+ R reachable, S stale,
+ D delay, P probe,
+ I incomplete, F failed,
+ N no-ARP, M
+ permanent (for example CR is configured and
+ reachable).
- unmap
- <iface>,<localAddr>[,<remoteAddr>]
+ unmap
+ <iface>,<localAddr>
+ [,<remoteAddr>]
Deletes matching tunnel endpoint or interface->address
associations previously set up with the map
- command.
+ command. A remote of dynamic stops neighbor
+ learning and removes learned remotes; explicit unicast
+ map entries are left in place.
@@ -1147,9 +1193,9 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]Gateway Commands
-
+
-
+
@@ -1166,8 +1212,8 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- rpush
- <srcIface>,<dstIfaceList>
+ rpush
+ <srcIface>,<dstIfaceList>
Resequence (modify IPv4 ID field or add IPv6 SMF-DPD option
header) and force relay of packets from
@@ -1222,100 +1268,818 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
GRE Tunnel Interfaces
- nrlsmf supports operation on Linux Generic
- Routing Encapsulation (GRE) interfaces, including point-to-point GRE and
- point-to-multipoint (mGRE) tunnels. GRE interfaces are treated
- differently from Ethernet interfaces: packets received on a GRE interface
- may not include a usable Ethernet header, and tunnel endpoint addresses
- (local and remote) are used for forwarding and elastic multicast control
- traffic.
-
- When a GRE interface is opened, nrlsmf
- attempts to read tunnel local and remote endpoint addresses from the
- operating system. If valid endpoint information is available, it is
- recorded automatically. Some GRE deployments (for example, lightweight
- or externally-managed "metadata" GRE tunnels) may not expose complete
- endpoint information to user space. In those cases, the
- "map" command must be used to associate tunnel
- endpoint addresses with the interface. A warning is logged when a GRE
- interface is used without adequate endpoint mapping.
-
- For mGRE operation, the remote endpoint may correspond to
- INADDR_ANY (0.0.0.0). The underlay network may
- require joining a multicast group so encapsulated traffic can be
- received; use the "ujoin" and
- "uleave" commands for this purpose on the physical
- underlay interface (not on the GRE interface itself).
-
- When elastic multicast is enabled, GRE tunnel endpoints are also
- used to identify the correct interface for elastic control messages
- (for example, EM_ACK). JSON configuration file support for GRE tunnel
- mapping is not yet available; use command-line or run-time remote
- control commands instead.
+ nrlsmf supports operation on Linux
+ Generic Routing Encapsulation (GRE) and point-to-multipoint (mGRE)
+ interfaces. From a configuration standpoint, a GRE or mGRE interface
+ is just another interface name that can be placed into an
+ nrlsmf interface group alongside physical
+ interfaces. Overlay unicast peer resolution (kernel
+ ip neigh, NHRP, or a multicast underlay group)
+ stays outside nrlsmf. Overlay
+ multicast inject onto a wildcard-remote mGRE
+ interface is the exception:
+ nrlsmf transmits once per mapped remote
+ (unicast peers and, optionally, an underlay multicast group).
+ Overlay GRE interfaces are typically
+ layered so a packet received on the tunnel is
+ not flooded back out of it. The NHRP hub is the exception: it must
+ replicate overlay multicast to the spokes. In every mode
+ nrlsmf still floods decapsulated overlay
+ multicast per its relay rules and keeps endpoint information for
+ duplicate packet detection (DPD) and Elastic Multicast.
+
+ GRE interfaces are treated differently from Ethernet interfaces:
+ packets received on a GRE interface may not include a usable Ethernet
+ header, and tunnel endpoint addresses (local and remote) are used for
+ forwarding and Elastic Multicast control traffic.
+
+
+ Endpoint Address Discovery
+
+ When a GRE interface is opened,
+ nrlsmf reads the tunnel's local and
+ remote endpoint addresses from the operating system. This is enough
+ for point-to-point GRE and for multicast-underlay mGRE (the
+ tunnel remote is already a multicast group): there is a single
+ remote, so one transmission has a well-defined destination.
+
+ map
+ <iface>,<localAddr>[,<remoteAddr>]
+ records tunnel endpoint information in
+ nrlsmf's interface-info table. Multiple
+ map commands with the same local address and
+ different remotes add multiple peers. That list is used in two
+ ways:
-
- GRE Tunnel Commands
+
+
+ DPD / Elastic Multicast bookkeeping (all GRE modes).
+
-
-
+
+ Overlay multicast inject onto a wildcard-remote
+ (0.0.0.0) mGRE interface: the kernel has
+ no single destination for overlay multicast, so
+ nrlsmf transmits once per mapped
+ remote that is a unicast peer or an underlay multicast group.
+ Unicast remotes come from explicit
+ map <iface>,<local>,<peer>
+ and/or
+ map <iface>,<local>,dynamic
+ (kernel neighbor table). A mapped remote may also be an
+ underlay multicast group; that is an explicit
+ map, not learned by
+ dynamic, and inbound GRE-in-multicast on
+ that group needs ujoin as well. Overlay
+ unicast on that same
+ interface still uses the kernel neighbor table
+ (ip neigh), not map.
+ map gre0,10.0.0.2,0.0.0.0 is allowed: it
+ records the kernel wildcard remote
+ (INADDR_ANY) for DPD / Elastic Multicast
+ bookkeeping. It is not a send destination and not a synonym
+ for dynamic. Opening a tunnel that already
+ has remote 0.0.0.0 stores the same
+ mapping, so that command is usually redundant.
+
+
-
+ The other case that requires map is a
+ "lightweight" or "metadata" GRE interface (externally-managed
+ tunnel with no fixed endpoints on the device).
+ nrlsmf logs a warning that the
+ interface is missing tunnel endpoint addressing until
+ map supplies it.
+
+ When Elastic Multicast is enabled, GRE tunnel endpoints are
+ also used to identify the correct interface for elastic control
+ messages (for example, EM_ACK). JSON configuration file support
+ for GRE tunnel mapping is not yet available; use command-line or
+ run-time remote control commands instead.
+
+
+
+ Point-to-Point GRE
+
+ A point-to-point GRE tunnel has one fixed local address and
+ one fixed remote address, configured once at tunnel setup. There
+ is no ambiguity about where an outbound packet goes, so no
+ ongoing peer-resolution mechanism is needed.
+
+ Topology:
+
+ Node A WAN / Underlay Node B
+ +---------------+ +---------------+
+ | eth0 (uplink)|<------------------------------------------->| eth0 (uplink)|
+ | 10.0.0.2 | | 10.0.1.2 |
+ | gre0 |=========== GRE tunnel (P2P) ================| gre0 |
+ | 172.16.0.1 | | 172.16.0.2 |
+ | wlan0 (MANET)| | eth1 (LAN) |
+ | SMF flooding | | multicast |
+ | | | receivers |
+ +---------------+ +---------------+
+
+ Kernel-level setup (Node A):
+
+ ip tunnel add gre0 mode gre local 10.0.0.2 remote 10.0.1.2 ttl 32
+ip addr add 172.16.0.1/30 dev gre0
+ip link set gre0 up
+
+ nrlsmf configuration (Node A):
+
+ nrlsmf resequence on \
+ cf wlan0 \
+ rpush wlan0,gre0 \
+ push gre0,wlan0 \
+ layered gre0
-
-
- map
- <iface>,<localAddr>[,<remoteAddr>]
-
- Associate tunnel endpoint address information with
- <iface>. If
- <remoteAddr> is omitted or invalid, the
- <localAddr> is recorded as an ordinary
- local interface address. If
- <remoteAddr> is valid, the mapping is
- recorded as tunnel endpoint information for GRE or other tunnel
- interfaces. For mGRE, <remoteAddr> may
- be 0.0.0.0.
-
+
+
+ cf wlan0 — classical flooding within
+ the MANET.
+
-
- unmap
- <iface>,<localAddr>[,<remoteAddr>]
+
+ rpush wlan0,gre0 — relay
+ MANET-originated traffic across the tunnel, tagging it for
+ DPD.
+
- Remove a previously configured tunnel or local address
- mapping from <iface>. The address
- arguments must match those supplied to
- "map".
-
+
+ push gre0,wlan0 — relay traffic from
+ the far side back into the MANET (already tagged).
+
-
- ujoin
- <groupAddr>,<iface>
+
+ layered gre0 — a point-to-point
+ tunnel gains nothing from being re-flooded out itself.
+
+
- Join multicast group <groupAddr>
- on underlay interface <iface> to enable
- reception of mGRE multicast-encapsulated traffic. The group
- address must be a valid IP multicast address.
-
+ Typical use case: connecting two MANET
+ "islands," or a MANET to a fixed-infrastructure multicast network,
+ across a WAN hop with no native multicast support.
+
+
+
+ Multipoint GRE (mGRE)
+
+ A single mGRE interface represents multiple remote peers
+ instead of one. GRE itself has no built-in signaling for which
+ peer a packet goes to; three different mechanisms answer that
+ question for overlay unicast, entirely below
+ nrlsmf.
+
+ Modes A and B are, at the kernel level, the same mechanism
+ — an mGRE interface with no fixed remote
+ (remote 0.0.0.0) and a neighbor-resolution
+ table that maps overlay tunnel addresses to underlay peer
+ addresses — differing only in whether that table is maintained by
+ a routing daemon or by hand. Mode C is fundamentally different:
+ instead of per-peer unicast resolution, it turns the tunnel into a
+ virtual broadcast/multicast segment.
+
+
+ Mode A — NHRP-resolved mGRE (dynamic hub-and-spoke)
+
+ Peers ("spokes") register their real underlay address with
+ a Next Hop Resolution Protocol (NHRP) server (typically a hub,
+ or "NHS" — Next Hop Server). Any peer needing to reach another
+ resolves the destination's tunnel-address-to-underlay-address
+ mapping dynamically through the NHS, and the kernel encapsulates
+ directly to the resolved address. This is the mechanism behind
+ Cisco DMVPN-style deployments; on Linux/FRR it is provided by
+ FRR's nhrpd daemon (RFC 2332), which
+ handles NHRP registration and resolution independently of
+ nrlsmf.
+ nrlsmf itself has no NHRP support
+ and does not need any: by the time a decapsulated packet
+ reaches nrlsmf on the mGRE
+ interface, nhrpd and the kernel have
+ already resolved and delivered it.
+
+ Topology:
+
+ Hub / NHRP Server (NHS)
+ +------------------------+
+ | eth0 (uplink) |
+ | 10.0.0.2 |
+ | gre0 (mGRE) |
+ | 172.16.0.1 |
+ +-----------+------------+
+ |
+ Underlay / WAN
+ dynamic NHRP resolution, per-peer unicast
+ / | \
+ +-------------------+ +-------------------+ +-------------------+
+ | Spoke S1 | | Spoke S2 | | Spoke S3 |
+ | eth0 10.0.1.2 | | eth0 10.0.2.2 | | eth0 10.0.3.2 |
+ | gre0 172.16.0.2 | | gre0 172.16.0.3 | | gre0 172.16.0.4 |
+ +-------------------+ +-------------------+ +-------------------+
+
+ Kernel + FRR setup (Hub):
+
+ ip tunnel add gre0 mode gre local 10.0.0.2 key 42 ttl 64
+ip addr add 172.16.0.1/32 dev gre0
+ip link set gre0 up
+
+ Omitting remote is the same as
+ remote 0.0.0.0 (kernel wildcard):
+
+ ip tunnel add gre0 mode gre local 10.0.0.2 remote 0.0.0.0 key 42 ttl 64
+
+ interface gre0
+ ip nhrp network-id 1
+ ip nhrp registration no-unique
+
+ The hub's own ordinary underlay address
+ (10.0.0.2, whatever interface it already
+ uses to reach the WAN) doubles as its NBMA identity — no
+ separate loopback or special addressing is required.
+
+ Kernel + FRR setup (Spoke S1; S2 and S3 use
+ 10.0.2.2/172.16.0.3 and
+ 10.0.3.2/172.16.0.4):
+
+ ip tunnel add gre0 mode gre local 10.0.1.2 key 42 ttl 64
+ip addr add 172.16.0.2/32 dev gre0
+ip link set gre0 up
+
+ Equivalent wildcard remote:
+
+ ip tunnel add gre0 mode gre local 10.0.1.2 remote 0.0.0.0 key 42 ttl 64
+
+ interface gre0
+ ip nhrp network-id 1
+ ip nhrp nhs 172.16.0.1 nbma 10.0.0.2
+ ip nhrp registration no-unique
+
+ nhrpd installs overlay-to-
+ underlay neighbors. A spoke's neighbor table is the NHS; the
+ hub's table is every registered spoke. Overlay unicast therefore
+ looks like static NBMA. Overlay multicast does not: a spoke has
+ only the hub as a mapped unicast remote, so overlay multicast
+ from a spoke goes to the hub, and the hub's
+ nrlsmf must replicate it to the
+ other spokes. Spokes are layered so a packet
+ received on the tunnel is not flooded back out of it. The hub
+ is not layered, because it is the overlay
+ replicator.
+
+ nrlsmf configuration (hub; each spoke
+ underlay address is a remote):
+
+ nrlsmf add overlay,cf,eth1,gre0 \
+ map gre0,10.0.0.2,10.0.1.2 \
+ map gre0,10.0.0.2,10.0.2.2 \
+ map gre0,10.0.0.2,10.0.3.2
+
+ Equivalent using neighbor learning
+ (nhrpd installs the same unicast
+ remotes in ip neigh):
+
+ nrlsmf add overlay,cf,eth1,gre0 \
+ map gre0,10.0.0.2,dynamic
+
+ nrlsmf configuration (spoke S1; S2 and S3 use
+ their own underlay address):
+
+ nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.1.2,dynamic
+
+ map …,dynamic is the usual inject
+ config with NHRP:
+ nhrpd programs
+ ip neigh, and
+ nrlsmf learns those unicast remotes.
+ Explicit maps of 0.0.0.0 record the kernel
+ wildcard remote for DPD / Elastic Multicast bookkeeping; they
+ are not send destinations and not a synonym for
+ dynamic.
+
+ Production DMVPN deployments typically add IPsec
+ (nhrpd integrates with strongSwan
+ via the VICI protocol) to protect the overlay traffic; that is
+ independent of nrlsmf and omitted
+ here. Spoke-to-spoke "shortcut" tunnels (DMVPN Phase 2/3)
+ require additional NFLOG/iptables configuration and are not
+ used for overlay multicast with
+ nrlsmf classic flooding — see FRR's
+ nhrpd documentation for that
+ setup.
+
+
+
+ Mode B — Statically-mapped mGRE (manual NBMA table)
+
+ Mechanically identical to Mode A at the kernel level —
+ same remote 0.0.0.0 mGRE interface — but
+ the peer-resolution table is populated by hand instead of by a
+ routing daemon.
+
+ Overlay unicast is resolved by the kernel neighbor table
+ (overlay-addr to underlay-addr). On Node A:
+
+ ip neigh add 172.16.0.2 lladdr 10.0.1.2 dev gre0 nud permanent
+ip neigh add 172.16.0.3 lladdr 10.0.2.2 dev gre0 nud permanent
+ip neigh add 172.16.0.4 lladdr 10.0.3.2 dev gre0 nud permanent
+
+ Overlay multicast has no entry in that table (and a
+ single send cannot fan out to several unicast underlay
+ destinations). nrlsmf injects
+ overlay multicast once per mapped remote in its
+ interface-info table. Those remotes are usually the same
+ unicast peers as in ip neigh, but a mapped
+ remote may also be an underlay multicast group (see
+ ).
+
+ The topology has the same shape as Mode A's diagram, with
+ static ip neigh entries in place of
+ nhrpd; typically used for a small,
+ fixed set of known peers.
+
+ nrlsmf configuration (Node A, local
+ underlay 10.0.0.2):
+
+ nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.0.2,10.0.1.2 \
+ map gre0,10.0.0.2,10.0.2.2 \
+ map gre0,10.0.0.2,10.0.3.2
+
+ Equivalent using neighbor learning from the
+ ip neigh table above
+ (dynamic dumps that table and keeps it
+ updated via NEWNEIGH/DELNEIGH):
+
+ nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.0.2,dynamic
+
+ The wildcard remote can be mapped explicitly. This is
+ allowed, but it is not a send destination and not a synonym
+ for dynamic. Opening
+ gre0 already records it from the kernel,
+ so it is usually redundant:
+
+ nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.0.2,0.0.0.0
+
+ Typical use case: a handful of
+ fixed, stable sites where the operational overhead of running
+ NHRP is not worth it, but a single mGRE interface (instead of
+ one point-to-point interface per peer) is still
+ convenient.
+
+
+
+ Mode C — Multicast-underlay mGRE
+
+ The tunnel's remote address is configured as an IP
+ multicast group address rather than a unicast peer or a
+ wildcard. A single encapsulated transmission is sent once, to
+ that group; the underlay network's own multicast routing (for
+ example, PIM) replicates and delivers it to every peer that
+ has joined the group. This is the mode
+ nrlsmf's
+ ujoin/uleave commands
+ exist for, and it is the most natural fit for SMF: symmetric,
+ no hub, no per-peer replication — the same job SMF already
+ does on a physical broadcast medium, relocated onto a
+ WAN.
+
+ Topology:
+
+ Underlay network with native IP multicast
+ routing (e.g. PIM), group 239.1.1.1
+ / | \
+ / | \
+ +-----------------+ +-----------------+ +-----------------+
+ | Node A | | Node B | | Node C |
+ | eth0 (underlay) | | eth0 (underlay) | | eth0 (underlay) |
+ | 10.0.0.2 | | 10.0.1.2 | | 10.0.2.2 |
+ | gre0 (mGRE) | | gre0 (mGRE) | | gre0 (mGRE) |
+ | 172.16.0.1 | | 172.16.0.2 | | 172.16.0.3 |
+ | remote=239.1.1.1| | remote=239.1.1.1| | remote=239.1.1.1|
+ | wlan0 (MANET) | | wlan0 (MANET) | | wlan0 (MANET) |
+ +-----------------+ +-----------------+ +-----------------+
+
+ Kernel-level setup (Node A; Node B and Node C
+ use 10.0.1.2/172.16.0.2
+ and
+ 10.0.2.2/172.16.0.3 —
+ symmetric, no hub role):
+
+ ip tunnel add gre0 mode gre local 10.0.0.2 remote 239.1.1.1 ttl 16
+ip addr add 172.16.0.1/24 dev gre0
+ip link set gre0 up
+
+ nrlsmf configuration (each node):
+
+ nrlsmf add overlay,cf,gre0,wlan0 \
+ layered gre0 \
+ ujoin 239.1.1.1,eth0
-
- uleave
- <groupAddr>,<iface>
+ ujoin/uleave are
+ issued against the underlay interface
+ (eth0), not the GRE interface itself — they
+ join or leave the multicast group on the physical network so
+ the kernel GRE device (whose remote is already that group)
+ actually receives the encapsulated traffic to decapsulate.
+ Mirror with uleave 239.1.1.1,eth0 on
+ shutdown or reconfiguration. This is not the same as mapping
+ a multicast group onto a wildcard-remote mGRE device; the
+ kernel already demuxes GRE-in-multicast onto the tunnel, so
+ nrlsmf does not also strip GRE from
+ underlay capture (that would double-process the packet).
+
+ When Elastic Multicast routing is enabled, the same
+ tunnel endpoint information is used to route EM control
+ traffic (for example, EM_ACK) to the
+ correct interface — another reason accurate endpoint state
+ matters even when data-plane decapsulation "just works."
+
+ This mode only works if the underlay genuinely supports
+ IP multicast routing end-to-end. If it does not, GRE-over-
+ multicast can silently fall back to resolving individual
+ neighbors instead of true multicast fan-out — worth verifying
+ the underlay's multicast path explicitly before relying on
+ this mode.
+
+
+
+ Mixed underlay: multicast group plus unicast peers
+
+ Wildcard-remote mGRE (Modes A and B) can inject overlay
+ multicast to an underlay multicast group as well as to unicast
+ peers. That is useful when some sites can join the group and
+ others cannot: one transmission covers the multicast-capable
+ set, and additional transmissions cover the unicast-only
+ peers. The kernel tunnel remote stays
+ 0.0.0.0; this is not multicast-underlay
+ mGRE (Mode C), whose device remote is already the group.
+
+ A Linux wildcard-remote mGRE device cannot send overlay
+ multicast as GRE-in-multicast by itself (there is no multicast
+ sll_addr for
+ PF_PACKET), and it does not demux inbound
+ GRE-in-multicast either: kernel GRE lookup matches the tunnel
+ local address, not the outer multicast destination. Mapping
+ the group as an inject remote makes
+ nrlsmf send GRE-in-multicast once
+ to that group. ujoin of the same group on
+ the underlay interface both joins IGMP and enables an underlay
+ capture so nrlsmf can strip GRE and
+ treat the inner packet as inbound on the overlay GRE
+ interface. Do not enable that capture on a multicast-underlay
+ GRE device: the kernel already delivers the packet on the
+ tunnel, and a second strip would double-process it.
+
+ Topology: five overlay routers on
+ one wildcard-remote mGRE. Nodes A, B, and C can join underlay
+ group 239.1.1.1. Nodes D and E cannot.
+
+ Node A -- lan0 --\
+ Node B -- lan1 ---\
+ Node C -- lan2 ---- underlay A/B/C join 239.1.1.1
+ Node D -- lan3 ---/ D/E unicast only
+ Node E -- lan4 --/
+
+ Kernel-level setup (Node A; B–E use their own
+ underlay and overlay addresses). Same wildcard-remote mGRE as
+ Modes A and B — the tunnel remote is
+ 0.0.0.0, not the underlay group:
+
+ ip tunnel add gre0 mode gre local 10.0.0.2 remote 0.0.0.0 ttl 64
+ip addr add 172.16.0.1/24 dev gre0
+ip link set gre0 multicast on
+ip link set gre0 up
+
+ Overlay unicast still needs kernel neighbors (Node A;
+ these can also be added by a dynamic protocol such as
+ NHRP):
+
+ ip neigh add 172.16.0.2 lladdr 10.0.1.2 dev gre0 nud permanent
+ip neigh add 172.16.0.3 lladdr 10.0.2.2 dev gre0 nud permanent
+ip neigh add 172.16.0.4 lladdr 10.0.3.2 dev gre0 nud permanent
+ip neigh add 172.16.0.5 lladdr 10.0.4.2 dev gre0 nud permanent
+
+ On A, B, and C, point the underlay group at the physical
+ iface so GRE-in-multicast has an output device:
+
+ ip route replace 239.1.1.1/32 dev eth0
+
+ nrlsmf configuration (Node A; B and C are the
+ same with their own underlay address):
+
+ nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ ujoin 239.1.1.1,eth0 \
+ map gre0,10.0.0.2,239.1.1.1 \
+ map gre0,10.0.0.2,10.0.3.2 \
+ map gre0,10.0.0.2,10.0.4.2
+
+ Overlay multicast inject from A is then three GRE
+ transmissions: one to 239.1.1.1, one to D
+ (10.0.3.2), and one to E
+ (10.0.4.2). Overlay unicast still uses
+ kernel ip neigh.
+
+ Nodes D and E create the same wildcard-remote tunnel
+ (Node D: local 10.0.3.2 / overlay
+ 172.16.0.4) and the same style of
+ ip neigh entries. They do not add the
+ underlay group route.
+
+ nrlsmf configuration (Node D; E is the same
+ with its own underlay address and no
+ ujoin):
+
+ nrlsmf add overlay,cf,eth1,gre0 \
+ layered gre0 \
+ map gre0,10.0.3.2,10.0.0.2 \
+ map gre0,10.0.3.2,10.0.1.2 \
+ map gre0,10.0.3.2,10.0.2.2 \
+ map gre0,10.0.3.2,10.0.4.2
+
+ layered on the overlay GRE interface
+ keeps a packet that arrived on the tunnel from being flooded
+ back out of it, so a unicast-only router does not re-inject
+ toward the multicast set. Traffic that arrives on
+ eth1 is still injected onto
+ gre0.
+
+
+
+
+ External (metadata) GRE
+
+ A fourth way GRE encapsulation parameters can be resolved,
+ distinct from all three mGRE modes above:
+ ip link add ... type gre external creates an
+ interface with no fixed local, remote, or key at all. Instead,
+ whatever adds routes or flow rules for it supplies the
+ encapsulation parameters per destination, using Linux's
+ lightweight tunnel ("lwtunnel") route encap:
+
+ ip route add 172.16.0.2/32 encap ip id 42 src 10.0.0.2 dst 10.0.1.2 ttl 16 dev gre0
+
+ This is the mechanism SDN controllers and OVS/OVN typically
+ use to build many per-flow or per-peer tunnels dynamically out of
+ a single device. It is conceptually the multipoint-resolution
+ equivalent of Mode B's static neighbor table — same job,
+ expressed as routes instead of neighbor entries — but because
+ the device itself has no fixed endpoint addresses, this is the
+ one case where nrlsmf's
+ map command is actually mandatory: there is
+ nothing for nrlsmf to auto-discover
+ from the interface, and it will log a warning
+ (GRE interface '<name>' missing tunnel endpoint
+ addressing. (must map it)) until map
+ supplies a local address. Overlay unicast still uses the kernel
+ lwtunnel routes. Overlay multicast inject has no single remote,
+ so nrlsmf also needs unicast remotes
+ in that same map table
+ (map <iface>,<local>,<peer>
+ per peer).
+ map …,0.0.0.0 is only the wildcard
+ bookkeeping form; map …,dynamic does not
+ apply (no ip neigh).
+
+ nrlsmf configuration (Node A):
+
+ nrlsmf add overlay,cf,gre0 \
+ layered gre0 \
+ map gre0,10.0.0.2,10.0.1.2 \
+ map gre0,10.0.0.2,10.0.2.2
+
+
+
+ Where Each Mode Fits
+
+
+ GRE / mGRE modes
+
+
+
+
+
+
+
+
+
+ Mode
+ Peer resolution
+ nrlsmf extras
+ Typical use
+
+
+
+
+
+ P2P GRE
+ Fixed at setup
+ typically layered
+ Two fixed sites; MANET-island bridging over a WAN
+ hop
+
+
+
+ mGRE Mode A (NHRP)
+ Dynamic, per-peer unicast (FRR
+ nhrpd)
+ map …,dynamic (or explicit
+ unicast maps); spokes
+ layered, hub not (replicator)
+ Hub-and-spoke overlay multicast; NHRP/DMVPN-style
+ tooling already in place
+
+
+
+ mGRE Mode B (static NBMA)
+ Static, per-peer unicast (manual
+ ip neigh)
+ explicit unicast maps (or
+ map …,dynamic); typically
+ layered
+ Small, fixed set of known peers; avoids running
+ NHRP
+
+
+
+ mGRE mixed underlay
+ Same wildcard-remote device as Modes A/B;
+ overlay multicast inject to an underlay group
+ and unicast peers
+ ujoin plus
+ map of the group (GRE-in-mcast send
+ and underlay demux); unicast
+ maps for peers that cannot join;
+ typically layered
+ Some sites can join underlay multicast, others
+ cannot
+
+
+
+ mGRE Mode C (multicast underlay)
+ None — underlay multicast fan-out (PIM)
+ ujoin/uleave
+ on the underlay interface (IGMP join; kernel demuxes
+ GRE-in-mcast onto the tunnel)
+ Symmetric SMF/MANET-gateway overlays
+
+
+
+ External (metadata) GRE
+ Per-destination lwtunnel routes (often
+ SDN-managed)
+ map is mandatory; remotes for
+ overlay multicast inject; typically
+ layered
+ SDN/OVS-managed overlays; an external controller
+ provisions tunnel endpoints
+
+
+
+
- Leave multicast group <groupAddr>
- on underlay interface <iface>.
-
-
-
-
+
+
+ Modes A and B are the same underlying mechanism — an
+ mGRE interface with no fixed remote and a peer-resolution
+ table — differing only in whether that table is maintained
+ dynamically (NHRP) or by hand. The same wildcard-remote
+ device can also map an underlay multicast group as one
+ overlay-multicast inject dest (mixed underlay); that is not
+ multicast-underlay mGRE (Mode C).
+
- Example: The following example adds GRE and
- Ethernet interfaces to a flooding group, maps explicit tunnel endpoints
- for a metadata GRE interface, and joins an underlay multicast group
- for mGRE reception:
+
+ Point-to-point and multicast-underlay mGRE (Mode C)
+ rely on nrlsmf's automatic
+ endpoint discovery (one configured remote). Modes A and B
+ overlay unicast use kernel neigh
+ (nhrpd or
+ ip neigh); overlay multicast inject uses
+ SMF's map table, filled by explicit
+ maps (unicast peers and/or an underlay
+ multicast group) and/or
+ map …,dynamic.
+
- nrlsmf add mygroup,cf,gre0,eth0 \
- map gre0,10.0.0.1,10.0.0.2 \
- ujoin 239.1.1.1,eth0
+
+ map is mandatory for
+ "external"/metadata GRE (no endpoints on the device). For
+ Mode A/B overlay multicast, mapped remotes (or
+ map …,dynamic)
+ are required. 0.0.0.0 means wildcard
+ remote, matching the kernel. It is not needed for
+ point-to-point or Mode C. Mapping an underlay multicast
+ group onto a wildcard-remote mGRE also requires
+ ujoin of that group so inbound
+ GRE-in-multicast can be demuxed.
+
+
+
+ In every mode, once packets reach
+ nrlsmf they are handled by the
+ same relay logic
+ (cf / smpr /
+ ecds / push /
+ rpush / merge /
+ rmerge / layered)
+ as any other interface. Overlay multicast inject onto a
+ wildcard-remote mGRE is the extra send-side step: one
+ transmission per mapped remote (unicast peer or underlay
+ group), still using that same relay decision. Overlay GRE
+ interfaces are typically layered; the
+ NHRP hub is not, because it replicates overlay multicast.
+
+
+
+
+
+ GRE Tunnel Commands
+
+
+ GRE Tunnel Commands
+
+
+
+
+
+
+
+
+ map
+ <iface>,<localAddr>
+ [,<remoteAddr>]
+
+ Associate tunnel endpoint address information with
+ <iface>. If
+ <remoteAddr> is omitted or
+ invalid, the <localAddr> is
+ recorded as an ordinary local interface address. If
+ <remoteAddr> is valid, the
+ mapping is recorded as tunnel endpoint information for
+ GRE or other tunnel interfaces. For mGRE,
+ <remoteAddr> may be 0.0.0.0
+ (kernel wildcard remote, not a send dest).
+ dynamic learns unicast remotes from
+ the kernel neighbor table (NHRP or static NBMA). It does
+ not apply to metadata/"external" GRE. Multiple commands
+ with different remotes record multiple inject
+ destinations (unicast peers and/or an underlay multicast
+ group) on a wildcard-remote mGRE interface or a metadata
+ GRE device.
+
+
+
+ unmap
+ <iface>,<localAddr>
+ [,<remoteAddr>]
+
+ Remove a previously configured tunnel or local
+ address mapping from
+ <iface>. The address arguments
+ must match those supplied to
+ "map". A remote of
+ dynamic stops neighbor learning and
+ removes learned remotes only.
+
+
+
+ ujoin
+ <groupAddr>,<iface>
+
+ Join multicast group
+ <groupAddr> on underlay
+ interface <iface> so the host
+ receives GRE-in-multicast. On multicast-underlay mGRE
+ this is an IGMP join so the kernel tunnel (remote
+ already the group) can decapsulate. On a wildcard-remote
+ mGRE device, if that same group is also a mapped inject
+ remote, this additionally enables underlay capture so
+ nrlsmf can strip GRE and
+ treat the inner packet as inbound on the overlay GRE
+ interface. The group address must be a valid IP
+ multicast address.
+
+
+
+ uleave
+ <groupAddr>,<iface>
+
+ Leave multicast group
+ <groupAddr> on underlay
+ interface <iface>.
+
+
+
+
+
@@ -1330,13 +2094,14 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]Virtual Interface Commands
-
-
+
+
- device
- <vifName>,<ifaceName>[/{t|r|d}][,<addr>...]
+ device
+ <vifName>,<ifaceName>
+ [/{t|r|d}][,<addr>...]
Create a virtual SMF interface (vif) associated with one
or more physical interfaces. Optional flags on
<ifaceName> are
@@ -1345,8 +2110,10 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- cid
- <vifName>,<iface1>[/{t|r|d}][,<iface2>...]
+ cid
+ <vifName>,
+ <iface1>[/{t|r|d}][,
+ <iface2>...]
Add or remove elements of a Composite Interface Device
associated with an SMF virtual interface.
@@ -1391,13 +2158,14 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]VRF Commands
-
-
+
+
- vrf
- <vrf-name>,[<vrf-id>,]<ifaceList>
+ vrf
+ <vrf-name>,
+ [<vrf-id>,]<ifaceList>
Associate one or more interfaces with the named VRF. An
optional numeric VRF identifier may be supplied.
@@ -1424,8 +2192,8 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]Configuration Commands
-
-
+
+
@@ -1468,14 +2236,14 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]Remote-only Commands
-
+
-
+
- selectorMac <binary
- macAddrArray>
+ selectorMac
+ <binary macAddrArray>
This command can be used by external processes (e.g.,
nrlolsrd) to control
@@ -1491,8 +2259,8 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- neighborMac <binary
- macAddrArray>
+ neighborMac
+ <binary macAddrArray>
This command can be used by external processes (e.g.,
nrlolsrd) to control
@@ -1508,8 +2276,8 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]
- mneBlockMac <binary
- macAddrArray>
+ mneBlockMac
+ <binary macAddrArray>
This command enables the
nrlsmf process to be compatible with
@@ -1618,14 +2386,103 @@ nrlsmf --cli [-i <instanceName>] <command> [args...]The same nrlsmf binary can act as a
client against that socket. Invoke it as
- "nrlsmf --cli [-i <instanceName>] <command>
- [args...]" to send a runtime command and, for status queries
- such as ping, stats,
- info, and the json* variants, print
- the reply. Configuration commands (for example
- "nrlsmf --cli debug 2") are sent without waiting for a
- response. Use "-i" when the target process was started
- with a non-default instance name.
+ "nrlsmf --cli [-i <instanceName>] -c "show <command>
+ [modifiers]"" to query runtime status. Repeat
+ -c to send more than one command (for example
+ "nrlsmf --cli -c "show statistics" -c "show interface grouping"").
+ Commands include
+ version, statistics,
+ interface, interface grouping
+ (a subcommand of interface),
+ tunnel, tunnel neighbors
+ (GRE/mGRE neighbors from configuration and/or the kernel),
+ groups, groups memberships,
+ and igmp groups. The last three require an Elastic
+ Multicast build. The unmodified command is the default listing.
+ Optional modifiers are command-specific. brief requests
+ less output and details requests more; a command may
+ support none, one, or both. When JSON is wanted, json
+ is always the last modifier (for example
+ "show groups brief json").
+ Older one-shot verbs such as stats and
+ groupsj still work for compatibility. Configuration
+ commands (for example "nrlsmf --cli -c "debug 2"",
+ "with-frr", or "elastic overlay")
+ are sent without waiting for a response. Use "-i" when
+ the target process was started with a non-default
+ instance name. Run "nrlsmf --cli ?"
+ for the show-command list.
+
+ On an Elastic build, show interface includes
+ flag M (JSON Managed) when
+ last-hop membership skip is in effect on that interface. The other
+ flags are L layered, T tunnel,
+ I IGMP proxy, and S
+ shadowing.
+
+ show groups lists detected multicast flows.
+ The fwd column is the flow default (often
+ L, LIMIT, about 1 packet per second).
+ Per-interface FORWARD after an EM_ACK is not that column; use
+ show groups details.
+ recv is the lifetime packet count.
+ pps is over a tumbling window of about 5
+ seconds.
+
+ nrlsmf --cli -i smf-r1 -c "show groups"
+Flags: A = Active, K = Ack
+Fwd: B = Block, H = Hybrid, L = Limit, F = Forward, D = Deny
+group saddr iif flags fwd recv pps
+---------------- ---------------- ------------ ----- --- ------ ----
+239.0.0.1 192.168.55.2 gre1 AK L 41 8.8
+
+ show groups details nests inbound sources
+ and outbound interfaces. Selected
+ Y is the current upstream.
+ Why is managed (IGMP or static
+ last hop), ack (EM_ACK from a downstream relay),
+ or -.
+
+ nrlsmf --cli -i smf-r1 -c "show groups details"
+Flags: A = Active, K = Ack
+Selected: Y = current upstream, N = additional inbound
+Fwd: B = Block, H = Hybrid, L = Limit, F = Forward, D = Deny
+Why: managed = IGMP/static last hop, ack = EM_ACK, - = default
+
+group saddr flags ack recv pps
+---------------- ---------------- ----- --- ------ ----
+239.0.0.1 192.168.55.2 AK yes 41 8.8
+ sources
+ iif upstream sel
+ gre1 10.0.0.2 Y
+ downstream
+ oif mode why fwd sent pps relays
+ eth1 elastic managed F 41 8.8 -
+ gre1 elastic ack F 0 0.0 -
+
+ show groups memberships is the controller
+ table: S static (join),
+ M managed (IGMP via with-frr),
+ E elastic (downstream EM_ACK). This is not the
+ last-hop skip list.
+
+ nrlsmf --cli -i smf-r1 -c "show groups memberships"
+Flags: S = Static, M = Managed, E = Elastic
+group saddr iface flags relays
+---------------- ---------------- ------------ ----- ----------------
+239.0.0.1 * eth1 M -
+239.0.0.1 192.168.55.2 gre1 E 10.0.0.2
+
+ show igmp groups is the last-hop list
+ filled by with-frr. A managed host interface that
+ does not list the group is not used as a last hop.
+
+ nrlsmf --cli -i smf-r1 -c "show igmp groups"
+Flags: M = Managed (last-hop HasActiveMembership)
+iface flags groups
+------------ ----- ------
+eth1 M 239.0.0.1
+gre1 - -
diff --git a/include/mcastFib.h b/include/mcastFib.h
index 05e4434..72277be 100644
--- a/include/mcastFib.h
+++ b/include/mcastFib.h
@@ -71,6 +71,39 @@ class MulticastFIB
void Update(unsigned int elapsedTime);
void Prune(unsigned int currentTime, unsigned int ageMax);
+ // Packet count plus a ~5s window for recent pps (ticks are microseconds).
+ class RateStat
+ {
+ public:
+ RateStat()
+ : total(0), window_count(0), window_start(0) {}
+ void Note(unsigned int tick)
+ {
+ total++;
+ if ((0 == window_start) || ((tick - window_start) > 5000000u))
+ {
+ window_start = tick;
+ window_count = 0;
+ }
+ window_count++;
+ }
+ unsigned int GetTotal() const
+ {return total;}
+ double GetPps(unsigned int now) const
+ {
+ if ((0 == window_start) || (0 == window_count))
+ return 0.0;
+ unsigned int dt = (now >= window_start) ? (now - window_start) : 0;
+ if (0 == dt)
+ return 0.0;
+ return (double)window_count * 1.0e6 / (double)dt;
+ }
+ private:
+ unsigned int total;
+ unsigned int window_count;
+ unsigned int window_start;
+ };
+
// TBD - what's the best way to sort our BucketList???
// a) time-ordered linked list like our Entry ActiveList, or
// b) indexed by interface index (using this for now)
@@ -89,6 +122,12 @@ class MulticastFIB
unsigned int GetInterfaceIndex() const
{return iface_index;}
+ void NoteSent(unsigned int tick)
+ {sent_stat.Note(tick);}
+ unsigned int GetSentPkts() const
+ {return sent_stat.GetTotal();}
+ double GetSentPps(unsigned int now) const
+ {return sent_stat.GetPps(now);}
void CopyStatus(const TokenBucket& bucket)
{
forwarding_status = bucket.forwarding_status;
@@ -120,6 +159,7 @@ class MulticastFIB
unsigned int token_interval; // microseconds per packet (1.0e+06 / packetsPerSecond)
unsigned int bucket_count;
unsigned int ticker_prev; // last time bucket was updated (microsecond ticks)
+ RateStat sent_stat;
}; // end class MulticastFIB::TokenBucket
// List of token buckets indexed by outbound iface index
@@ -700,6 +740,12 @@ class MulticastFIB
unsigned int Age(unsigned int currentTick);
unsigned int GetUpdateCount() const
{return update_count;}
+ void NoteRecv(unsigned int tick)
+ {recv_stat.Note(tick);}
+ unsigned int GetRecvPkts() const
+ {return recv_stat.GetTotal();}
+ double GetRecvPps(unsigned int now) const
+ {return recv_stat.GetPps(now);}
unsigned int GetUpdateInterval() const;
bool UpdatePending() const;
//unsigned int GetAge(unsigned int currentTick) const;
@@ -762,6 +808,7 @@ class MulticastFIB
unsigned int acking_interval_max; // in microseconds
unsigned int acking_interval_min; // in microseconds
UINT8 flow_ttl;
+ RateStat recv_stat;
Entry* active_prev;
Entry* active_next;
@@ -937,6 +984,8 @@ class MulticastFIB
{return downstream_relay_count;}
void PrintDownstreamRelayList(FILE* filePtr = NULL); // to ProtoDebug by default
+ DownstreamRelayList& AccessDownstreamRelayList()
+ {return downstream_relay_list;}
/*void SetUpstreamRelayAddress(const ProtoAddress& relayAddr, const ProtoAddress& advAddr)
{
@@ -1223,8 +1272,10 @@ class MulticastFIB
#ifdef ADAPTIVE_ROUTING
bool ParseFlowList( ProtoPktIP& pkt, Entry*& fibEntry, unsigned int currentTick, bool& sendAck,const ProtoAddress& srcMac);
#endif // ADAPTIVE_ROUTING
- void DumpFlowList(bool brief, std::ostringstream& ss);
- void DumpFlowListJson(bool brief, std::ostringstream& ss);
+ void DumpFlowList(bool brief, std::ostringstream& ss, bool details = false,
+ unsigned int currentTick = 0);
+ void DumpFlowListJson(bool brief, std::ostringstream& ss, bool details = false,
+ unsigned int currentTick = 0);
private:
EntryTable flow_table; // Table of detected flows (updated by forwarding plane)
@@ -1352,11 +1403,16 @@ class ElasticMulticastForwarder
{
public:
virtual bool SendFrame(unsigned int ifaceIndex, char* buffer, unsigned int length) = 0;
+ // Optional GRE dest (mGRE EM_ACK to one underlay peer).
+ virtual bool SendFrameTo(unsigned int ifaceIndex, char* buffer, unsigned int length,
+ const ProtoAddress& dest)
+ {return SendFrame(ifaceIndex, buffer, length);}
}; // end class ElasticMulticastForwarder::OutputMechanism
void SetOutputMechanism(OutputMechanism* mech)
{output_mechanism = mech;}
- void DumpGroups(bool brief, bool useJson, std::ostringstream& ss);
+ void DumpGroups(bool brief, bool useJson, std::ostringstream& ss, bool details = false);
+ void DumpManagedGroups(bool useJson, std::ostringstream& ss);
protected:
// Our "ticker" is a count of microseconds that is used for our
@@ -1450,7 +1506,9 @@ class ElasticMulticastController
MulticastFIB::MembershipTable& AccessMembershipTable()
{return membership_table;}
- void DumpGroups(bool brief, bool useJson, std::ostringstream& ss);
+ void DumpGroups(bool brief, bool useJson, std::ostringstream& ss, bool details = false);
+ void DumpMemberships(bool useJson, std::ostringstream& ss);
+ void DumpManagedGroups(bool useJson, std::ostringstream& ss);
// NEXT STEP - IMPLEMENT MECHANISM TO SEND ACKS to UPSTREAM FORWARDERS
// 1) When do we send an ACK?
diff --git a/include/smf.h b/include/smf.h
index 99df3ae..882a419 100644
--- a/include/smf.h
+++ b/include/smf.h
@@ -109,14 +109,17 @@ class Smf
InterfaceInfo(unsigned int ifaceIndex,
const ProtoAddress& localAddr,
const ProtoAddress* remoteAddr = NULL,
- bool mapped = false)
- : iface_index(ifaceIndex), local_addr(localAddr), is_mapped(mapped)
+ bool fromConfig = false,
+ bool fromKernel = false)
+ : iface_index(ifaceIndex), local_addr(localAddr),
+ from_config(fromConfig), from_kernel(fromKernel), is_learned(false)
{
// address_info_key is tuple of [remoteAddr]localAddr (i.e. remoteAddr is optional)
unsigned int len = 0;
if (NULL != remoteAddr)
{
// remoteAddr is first in key for FindTunnelInfo() for mGRE to work
+ remote_addr = *remoteAddr;
len = remoteAddr->GetLength();
memcpy(address_info_key, remoteAddr->GetRawHostAddress(), len);
}
@@ -128,10 +131,18 @@ class Smf
~InterfaceInfo() {}
void SetIndex(unsigned int index) {iface_index = index;}
void SetMaskLength(unsigned int maskLen) {local_mask_len = maskLen;}
+ void MarkConfig() {from_config = true;}
+ void MarkKernel() {from_kernel = true;}
unsigned int GetIndex() const {return iface_index;}
unsigned int GetMaskLength() const {return local_mask_len;}
const ProtoAddress& GetLocalAddress() const {return local_addr;}
const ProtoAddress& GetRemoteAddress() const {return remote_addr;}
+ void SetMapped(bool state) {from_config = state;}
+ bool IsMapped() const {return from_config;}
+ void SetLearned(bool state) {is_learned = state;}
+ bool IsLearned() const {return is_learned;}
+ bool FromConfig() const {return from_config;}
+ bool FromKernel() const {return from_kernel;}
private:
// Required ProtoTreeItem overrides
@@ -142,8 +153,9 @@ class Smf
ProtoAddress local_addr;
unsigned int local_mask_len;
ProtoAddress remote_addr; // invalid for non-tunnels, INADDR_ANY for mGRE tunnels
- bool is_mapped; // false for interfaces assigned to the interface, true for "mapped" association
- // (stored but not yet used for lookup or policy)
+ bool from_config; // true if added with the map command
+ bool from_kernel; // true if learned from GRE device attributes
+ bool is_learned; // true for kernel-neigh learned remotes (map ...,dynamic)
char address_info_key[16+16]; // big enough for IPv6
unsigned int address_info_size; // in bits
}; // end class Smf::InterfaceInfo
@@ -154,12 +166,13 @@ class Smf
InterfaceInfo* InsertIndex(unsigned int ifaceIndex,
const ProtoAddress& localAddr,
const ProtoAddress* remoteAddr = NULL,
- bool mapped = false)
+ bool fromConfig = false,
+ bool fromKernel = false)
{
InterfaceInfo* info = (NULL == remoteAddr) ? FindInfo(localAddr) : FindInfo(localAddr, *remoteAddr);
if (NULL == info)
{
- info = new InterfaceInfo(ifaceIndex, localAddr, remoteAddr, mapped);
+ info = new InterfaceInfo(ifaceIndex, localAddr, remoteAddr, fromConfig, fromKernel);
if (NULL == info)
{
PLOG(PL_ERROR, "InterfaceInfoTable::InsertIndex() new InterfaceInfo error: %s\n", GetErrorString());
@@ -174,6 +187,8 @@ class Smf
else
{
info->SetIndex(ifaceIndex);
+ if (fromConfig) info->MarkConfig();
+ if (fromKernel) info->MarkKernel();
}
return info;
}
@@ -181,10 +196,12 @@ class Smf
{return Find(addr.GetRawHostAddress(), 8*addr.GetLength());}
InterfaceInfo* FindInfo(const ProtoAddress& localAddr, const ProtoAddress& remoteAddr) const
{
+ // Must match InterfaceInfo key order: [remoteAddr][localAddr]
char addrInfo[16 + 16];
- memcpy(addrInfo, localAddr.GetRawHostAddress(), localAddr.GetLength());
- memcpy(addrInfo+localAddr.GetLength(), remoteAddr.GetRawHostAddress(), remoteAddr.GetLength());
- return Find(addrInfo, 8*(localAddr.GetLength() + remoteAddr.GetLength()));
+ unsigned int len = remoteAddr.GetLength();
+ memcpy(addrInfo, remoteAddr.GetRawHostAddress(), len);
+ memcpy(addrInfo + len, localAddr.GetRawHostAddress(), localAddr.GetLength());
+ return Find(addrInfo, 8*(len + localAddr.GetLength()));
}
unsigned int GetIndex(const ProtoAddress& addr) const
{
@@ -273,28 +290,48 @@ class Smf
unsigned int GetInterfaceIndex(const ProtoAddress& addr) const
{return iface_info_table.GetIndex(addr);}
- bool AddTunnelInfo(unsigned int ifaceIndex, const ProtoAddress& localAddr, const ProtoAddress& remoteAddr)
+ bool AddTunnelInfo(unsigned int ifaceIndex, const ProtoAddress& localAddr, const ProtoAddress& remoteAddr,
+ bool mapped = true, bool learned = false)
{
TRACE("mapping tunnel addrs local:%s", localAddr.GetHostString());
TRACE(" remote:%s\n", remoteAddr.GetHostString());
- InterfaceInfo* ifaceInfo = iface_info_table.InsertIndex(ifaceIndex, localAddr, &remoteAddr, true);
+ bool fromKernel = !mapped && !learned;
+ InterfaceInfo* ifaceInfo = iface_info_table.InsertIndex(ifaceIndex, localAddr, &remoteAddr,
+ mapped, fromKernel);
if (NULL == ifaceInfo)
{
PLOG(PL_ERROR, "Smf::AddTunnelInfo() iface_info_table.InsertIndex() failed\n");
return false;
}
+ if (mapped) ifaceInfo->SetMapped(true);
+ if (learned) ifaceInfo->SetLearned(true);
unsigned int maskLen = ProtoNet::GetInterfaceAddressMask(ifaceIndex, localAddr);
ifaceInfo->SetMaskLength(maskLen);
return true;
}
void RemoveTunnelInfo(const ProtoAddress& localAddr, const ProtoAddress& remoteAddr)
- {iface_info_table.RemoveAddress(localAddr, &remoteAddr);}
+ {ClearTunnelSource(localAddr, remoteAddr, true, true);}
+ void ClearTunnelSource(const ProtoAddress& localAddr, const ProtoAddress& remoteAddr,
+ bool mapped, bool learned)
+ {
+ InterfaceInfo* info = iface_info_table.FindInfo(localAddr, remoteAddr);
+ if (NULL == info)
+ return;
+ if (mapped) info->SetMapped(false);
+ if (learned) info->SetLearned(false);
+ if (!info->IsMapped() && !info->IsLearned())
+ iface_info_table.RemoveAddress(localAddr, &remoteAddr);
+ }
unsigned int GetTunnelIndex(const ProtoAddress& localAddr, const ProtoAddress& remoteAddr) const
{
// For point-to-point GRE tunnel interfaces where endpoint information is explicit
return iface_info_table.GetIndex(localAddr, remoteAddr);
}
+ // Match an EM_ACK upstream addr to *this* node's tunnel local or
+ // overlay IP. Do not use GetInterfaceIndex() here: map remotes are
+ // also in that table, so a peer underlay would match every spoke.
+ unsigned int FindInterfaceByLocalEndpoint(const ProtoAddress& addr);
unsigned int FindTunnelIndex(const ProtoAddress& localAddr, const ProtoAddress& remoteAddr)
{
// For point-to-multipoint (mGRE) tunnel interfaces.
@@ -332,6 +369,62 @@ class Smf
InterfaceInfoTable& AccessInterfaceInfoTable()
{return iface_info_table;}
+ // Mapped remotes used as GRE inject destinations for overlay
+ // multicast: unicast peers and (optionally) an underlay multicast
+ // group. Skip 0.0.0.0 (kernel wildcard, not a send dest).
+ void GetTunnelUnicastRemotes(unsigned int ifaceIndex, ProtoAddressList& dests)
+ {
+ InterfaceInfoTable::Iterator iterator(iface_info_table);
+ InterfaceInfo* info;
+ while (NULL != (info = iterator.GetNextItem()))
+ {
+ if (info->GetIndex() != ifaceIndex)
+ continue;
+ const ProtoAddress& remote = info->GetRemoteAddress();
+ if (remote.IsValid() && (remote.IsUnicast() || remote.IsMulticast()))
+ dests.Insert(remote);
+ }
+ }
+ // First mapped underlay multicast remote (multicast-underlay mGRE).
+ bool GetTunnelMulticastRemote(unsigned int ifaceIndex, ProtoAddress& dest)
+ {
+ InterfaceInfoTable::Iterator iterator(iface_info_table);
+ InterfaceInfo* info;
+ while (NULL != (info = iterator.GetNextItem()))
+ {
+ if (info->GetIndex() != ifaceIndex)
+ continue;
+ const ProtoAddress& remote = info->GetRemoteAddress();
+ if (remote.IsValid() && remote.IsMulticast())
+ {
+ dest = remote;
+ return true;
+ }
+ }
+ return false;
+ }
+ // True if addr is a unicast GRE peer on this iface (map or learned).
+ bool FindTunnelUnicastPeer(unsigned int ifaceIndex, const ProtoAddress& addr,
+ ProtoAddress& dest);
+ bool FindOverlayForUnderlay(unsigned int ifaceIndex, const ProtoAddress& underlay,
+ ProtoAddress& overlay);
+
+ unsigned int FindMappedIndexForRemote(const ProtoAddress& remote)
+ {
+ if (!remote.IsValid())
+ return 0;
+ InterfaceInfoTable::Iterator iterator(iface_info_table);
+ InterfaceInfo* info;
+ while (NULL != (info = iterator.GetNextItem()))
+ {
+ if (info->IsMapped() &&
+ info->GetRemoteAddress().IsValid() &&
+ info->GetRemoteAddress().HostIsEqual(remote))
+ return info->GetIndex();
+ }
+ return 0;
+ }
+
UINT16 GetIPv4LocalSequence(const ProtoAddress* dstAddr,
const ProtoAddress* srcAddr = NULL)
{return ((UINT16)ip4_seq_mgr.GetSequence(dstAddr, srcAddr));}
@@ -396,6 +489,13 @@ class Smf
{tunnel_remote_addr = addr;}
const ProtoAddress& GetTunnelRemoteAddress() const
{return tunnel_remote_addr;}
+ void SetTunnelLearnDynamic(bool state)
+ {tunnel_learn_dynamic = state;}
+ bool GetTunnelLearnDynamic() const
+ {return tunnel_learn_dynamic;}
+ ProtoAddressList& AccessLearnedOverlays()
+ {return learned_overlays;}
+ void ClearLearnedOverlays();
bool IsGRE() const
{return tunnel_local_addr.IsValid();}
@@ -563,6 +663,8 @@ class Smf
{managed_memberships.Remove(grpAddr);}
bool HasActiveMembership(const ProtoAddress& grpAddr) const
{return managed_memberships.Contains(grpAddr);}
+ ProtoAddressList& AccessManagedMemberships()
+ {return managed_memberships;}
#endif // ELASTIC_MCAST
// This is for adding an opaque "decorator" extension to the interface
@@ -669,6 +771,8 @@ class Smf
ProtoAddressList addr_list; // list of IP addresses of the interface
ProtoAddress tunnel_local_addr; // valid when Smf::Interface is GRE endpoint
ProtoAddress tunnel_remote_addr;
+ bool tunnel_learn_dynamic; // map ,,dynamic
+ ProtoAddressList learned_overlays; // overlay neigh dst -> underlay (userData)
ProtoAddress ip_addr; // used as source addr for nrlsmf IPIP encapsulation
std::string if_name;
bool resequence;
diff --git a/include/smfIgmp.h b/include/smfIgmp.h
index 4fcc908..e744d2e 100644
--- a/include/smfIgmp.h
+++ b/include/smfIgmp.h
@@ -38,6 +38,9 @@ class SmfIgmp : public ProtoChannel
virtual bool Open(bool withFRR);
virtual void Close();
virtual bool IsOpen() const;
+ // Start (or resume) polling FRR PIM for IGMP ifaces/groups.
+ // Safe after Open(false) so "with-frr" can be sent at runtime.
+ void EnableFrrPolling();
void ProcessUpdates();
bool HasMembershipUpdates() const
diff --git a/include/smfVersion.h b/include/smfVersion.h
index ef9d3c1..95da114 100644
--- a/include/smfVersion.h
+++ b/include/smfVersion.h
@@ -1,3 +1,3 @@
#ifndef _SMF_VERSION
-#define _SMF_VERSION "1.3"
+#define _SMF_VERSION "1.4"
#endif // !_SMF_VERSION
diff --git a/protolib b/protolib
index 73dc51a..00e09e7 160000
--- a/protolib
+++ b/protolib
@@ -1 +1 @@
-Subproject commit 73dc51a2479a355b7cc05e34ec81f089588f4b50
+Subproject commit 00e09e7a1f76142cc707c2f78cb16e0e1492aa94
diff --git a/src/common/mcastFib.cpp b/src/common/mcastFib.cpp
index e993ba0..7b04fed 100644
--- a/src/common/mcastFib.cpp
+++ b/src/common/mcastFib.cpp
@@ -1989,49 +1989,215 @@ void MulticastFIB::PruneFlowList(unsigned int currentTick, ElasticMulticastContr
}
} // end MulticastFIB::PruneFlowList()
-void MulticastFIB::DumpFlowList(bool brief, std::ostringstream& ss)
+static void FormatIfaceName(unsigned int ifaceIndex, char* ifaceName, unsigned int nameMax,
+ const char* missing = "")
+{
+ strncpy(ifaceName, missing, nameMax);
+ ifaceName[nameMax] = '\0';
+ if (0 != ifaceIndex)
+ ProtoNet::GetInterfaceName(ifaceIndex, ifaceName, nameMax);
+}
+
+static char FwdStatusLetter(MulticastFIB::ForwardingStatus status)
+{
+ switch (status)
+ {
+ case MulticastFIB::BLOCK: return 'B';
+ case MulticastFIB::HYBRID: return 'H';
+ case MulticastFIB::LIMIT: return 'L';
+ case MulticastFIB::FORWARD: return 'F';
+ case MulticastFIB::DENY: return 'D';
+ default: return '?';
+ }
+}
+
+static void FormatFlowFlagLetters(bool active, bool ack, char* buf, size_t bufLen)
+{
+ size_t n = 0;
+ if (active && (n + 1 < bufLen)) buf[n++] = 'A';
+ if (ack && (n + 1 < bufLen)) buf[n++] = 'K';
+ if ((0 == n) && (n + 1 < bufLen)) buf[n++] = '-';
+ if (bufLen > 0) buf[n] = '\0';
+}
+
+static void FormatMembershipFlagLetters(int flags, char* buf, size_t bufLen)
+{
+ size_t n = 0;
+ if ((0 != (flags & MulticastFIB::Membership::STATIC)) && (n + 1 < bufLen))
+ buf[n++] = 'S';
+ if ((0 != (flags & MulticastFIB::Membership::MANAGED)) && (n + 1 < bufLen))
+ buf[n++] = 'M';
+ if ((0 != (flags & MulticastFIB::Membership::ELASTIC)) && (n + 1 < bufLen))
+ buf[n++] = 'E';
+ if ((0 == n) && (n + 1 < bufLen)) buf[n++] = '-';
+ if (bufLen > 0) buf[n] = '\0';
+}
+
+static void DumpOutboundBuckets(MulticastFIB::Entry& entry, std::ostringstream& ss, bool json,
+ unsigned int currentTick)
+{
+ MulticastFIB::BucketList::Iterator iterator(entry.AccessBucketList());
+ MulticastFIB::TokenBucket* bucket;
+ bool comma = false;
+ if (json)
+ ss << "[";
+ while (NULL != (bucket = iterator.GetNextItem()))
+ {
+ char ifaceName[Smf::IF_NAME_MAX + 1];
+ FormatIfaceName(bucket->GetInterfaceIndex(), ifaceName, Smf::IF_NAME_MAX);
+ const char* status = MulticastFIB::GetForwardingStatusString(bucket->GetForwardingStatus());
+ if (json)
+ {
+ ss << (comma ? "," : "")
+ << "{\"Interface\" : \"" << ifaceName
+ << "\", \"FwdStatus\" : \"" << status
+ << "\", \"SentPkts\" : " << bucket->GetSentPkts()
+ << ", \"SentPps\" : " << std::fixed << std::setprecision(1)
+ << bucket->GetSentPps(currentTick) << "}";
+ }
+ else
+ {
+ ss << (comma ? " " : "") << ifaceName << ":"
+ << FwdStatusLetter(bucket->GetForwardingStatus())
+ << "(" << bucket->GetSentPkts() << ","
+ << std::fixed << std::setprecision(1)
+ << bucket->GetSentPps(currentTick) << ")";
+ }
+ comma = true;
+ }
+ if (json)
+ ss << "]";
+ else if (!comma)
+ ss << "-";
+}
+
+static void DumpMembershipFlags(int flags, std::ostringstream& ss)
+{
+ bool first = true;
+ ss << "[";
+ if (0 != (flags & MulticastFIB::Membership::STATIC))
+ {
+ ss << "\"STATIC\"";
+ first = false;
+ }
+ if (0 != (flags & MulticastFIB::Membership::MANAGED))
+ {
+ if (!first)
+ ss << ",";
+ ss << "\"MANAGED\"";
+ first = false;
+ }
+ if (0 != (flags & MulticastFIB::Membership::ELASTIC))
+ {
+ if (!first)
+ ss << ",";
+ ss << "\"ELASTIC\"";
+ first = false;
+ }
+ ss << "]";
+}
+
+static void DumpDownstreamRelays(MulticastFIB::Membership& membership, std::ostringstream& ss, bool json)
+{
+ MulticastFIB::DownstreamRelayList::Iterator iterator(membership.AccessDownstreamRelayList());
+ MulticastFIB::DownstreamRelay* relay;
+ bool comma = false;
+ if (json)
+ ss << "[";
+ while (NULL != (relay = iterator.GetNextItem()))
+ {
+ char host[256] = "";
+ relay->GetIpAddr().GetHostString(host, 255);
+ if (json)
+ ss << (comma ? "," : "") << "\"" << host << "\"";
+ else
+ ss << (comma ? "," : "") << host;
+ comma = true;
+ }
+ if (json)
+ ss << "]";
+ else if (!comma)
+ ss << "-";
+}
+
+void MulticastFIB::DumpFlowList(bool brief, std::ostringstream& ss, bool details,
+ unsigned int currentTick)
{
// pass over the flow_table and dump text suitable for command output
MulticastFIB::Entry* entry = NULL;
MulticastFIB::EntryTable::Iterator fiberator(flow_table);
- if (brief) {
- ss << "MCast Address SRC Interface\n";
- ss << "---------------- -------------\n";
- } else {
- ss << "MCast Address SRC Address Status Fwd Status ACK SRC Interface\n";
- ss << "---------------- ---------------- ------ ---------- --- -------------\n";
+ if (brief)
+ {
+ ss << "group iif\n"
+ << "---------------- ------------\n";
+ }
+ else
+ {
+ ss << "Flags: A = Active, K = Ack\n"
+ << "Fwd: B = Block, H = Hybrid, L = Limit, F = Forward, D = Deny\n";
+ if (details)
+ {
+ ss << "group saddr upstream iif flags fwd recv pps oif\n"
+ << "---------------- ---------------- ---------------- ------------ ----- --- ------ ---- ----\n";
+ }
+ else
+ {
+ ss << "group saddr iif flags fwd recv pps\n"
+ << "---------------- ---------------- ------------ ----- --- ------ ----\n";
+ }
}
- while (NULL != (entry = fiberator.GetNextEntry()))
+ while (NULL != (entry = fiberator.GetNextEntry()))
{
- ProtoFlow::Description flow;
ProtoAddress dst, src;
- char ifaceName[Smf::IF_NAME_MAX+1] = "";
- UpstreamRelay *up;
+ char group[64] = "-";
+ char saddr[64] = "*";
+ char uaddr[64] = "-";
+ char iif[Smf::IF_NAME_MAX + 1];
+ char flags[8];
+ UpstreamRelay* up = entry->GetCurrentBestUpstreamRelay();
- flow = entry->GetFlowDescription();
- up = entry->GetCurrentBestUpstreamRelay();
entry->GetDstAddr(dst);
- ss << std::left << std::setw(17) << dst.GetHostString();
- if (!brief) {
- char srcHostString[256] = "*"; // initialize this, as GetHostString() is a bad actor
-
+ dst.GetHostString(group, sizeof(group) - 1);
+ if (!brief)
+ {
entry->GetSrcAddr(src);
- src.GetHostString(srcHostString,255);
- ss << std::left << std::setw(17) << srcHostString;
- ss << (entry->IsActive() ? "ACTIVE " : "IDLE ");
- ss << std::left << std::setw(11) << MulticastFIB::GetForwardingStatusString(entry->GetDefaultForwardingStatus());
- ss << (entry->GetAckingStatus() ? " x " : " ");
+ if (src.IsValid())
+ src.GetHostString(saddr, sizeof(saddr) - 1);
+ FormatFlowFlagLetters(entry->IsActive(), entry->GetAckingStatus(), flags, sizeof(flags));
+ }
+ FormatIfaceName(up ? up->GetInterfaceIndex() : 0, iif, Smf::IF_NAME_MAX, "-");
+ if (details && up)
+ up->GetAddress().GetHostString(uaddr, sizeof(uaddr) - 1);
+
+ ss << std::left << std::setw(16) << group << " ";
+ if (brief)
+ {
+ ss << iif << "\n";
+ continue;
+ }
+ ss << std::setw(16) << saddr << " ";
+ if (details)
+ ss << std::setw(16) << uaddr << " ";
+ ss << std::setw(12) << iif << " ";
+ ss << std::setw(5) << flags << " ";
+ ss << FwdStatusLetter(entry->GetDefaultForwardingStatus());
+ ss << " " << std::setw(6) << entry->GetRecvPkts() << " ";
+ ss << std::fixed << std::setprecision(1) << std::setw(4)
+ << entry->GetRecvPps(currentTick);
+ if (details)
+ {
+ ss << " ";
+ DumpOutboundBuckets(*entry, ss, false, currentTick);
}
- if (up) ProtoNet::GetInterfaceName(up->GetInterfaceIndex(), ifaceName, Smf::IF_NAME_MAX);
- ss << ifaceName;
ss << "\n";
}
} // end MulticastFIB::DumpFlowList()
-void MulticastFIB::DumpFlowListJson(bool brief, std::ostringstream& ss)
+void MulticastFIB::DumpFlowListJson(bool brief, std::ostringstream& ss, bool details,
+ unsigned int currentTick)
{
// pass over the flow_table and dump text suitable for json output
MulticastFIB::Entry* entry = NULL;
@@ -2061,9 +2227,21 @@ void MulticastFIB::DumpFlowListJson(bool brief, std::ostringstream& ss)
ss << "\"Status\" : \"" << (entry->IsActive() ? "ACTIVE" : "IDLE") << "\",";
ss << "\"FwdStatus\" : \"" << MulticastFIB::GetForwardingStatusString(entry->GetDefaultForwardingStatus()) << "\",";
ss << "\"Ack\" : \"" << (entry->GetAckingStatus() ? "yes" : "no") << "\",";
+ ss << "\"RecvPkts\" : " << entry->GetRecvPkts() << ",";
+ ss << "\"RecvPps\" : " << std::fixed << std::setprecision(1)
+ << entry->GetRecvPps(currentTick) << ",";
}
if (up) ProtoNet::GetInterfaceName(up->GetInterfaceIndex(), ifaceName, Smf::IF_NAME_MAX);
ss << "\"SrcInterface\" : \"" << ifaceName << "\"";
+ if (details)
+ {
+ char upHost[256] = "";
+ if (up)
+ up->GetAddress().GetHostString(upHost, 255);
+ ss << ", \"UpstreamAddr\" : \"" << upHost << "\"";
+ ss << ", \"Outbound\" : ";
+ DumpOutboundBuckets(*entry, ss, true, currentTick);
+ }
ss << "}";
comma = true;
}
@@ -2222,8 +2400,12 @@ bool ElasticMulticastForwarder::SetForwardingStatus(const ProtoFlow::Description
// static forwarding entries).
// Sets forwarding status for all matching (sub-matching, non-bimatch) entries.
+ // IGMP memberships are dst-only (*,G). FLAG_ALL on that description
+ // never hits the live (S,G) packet FIB entry.
+ int matchFlags = (0 == flowDescription.GetSrcLength()) ?
+ ProtoFlow::Description::FLAG_DST : ProtoFlow::Description::FLAG_ALL;
MulticastFIB::EntryTable& flowTable = mcast_fib.AccessFlowTable();
- MulticastFIB::EntryTable::Iterator iterator(flowTable, &flowDescription, ProtoFlow::Description::FLAG_ALL, false); // non bimatch iterator
+ MulticastFIB::EntryTable::Iterator iterator(flowTable, &flowDescription, matchFlags, false); // non bimatch iterator
MulticastFIB::Entry* entry = iterator.GetNextEntry();
if (NULL != entry)
{
@@ -2302,10 +2484,283 @@ bool ElasticMulticastForwarder::SetManagedStatus(const ProtoFlow::Description& f
return true;
} // end ElasticMulticastForwarder::SetManagedStatus()
-void ElasticMulticastForwarder::DumpGroups(bool brief, bool useJson, std::ostringstream& ss)
+static MulticastFIB::Membership* FindIfaceMembership(
+ MulticastFIB::MembershipTable& table,
+ unsigned int ifaceIndex,
+ const ProtoAddress& group)
+{
+ MulticastFIB::MembershipTable::Iterator iterator(table);
+ MulticastFIB::Membership* membership;
+ while (NULL != (membership = iterator.GetNextEntry()))
+ {
+ if (membership->GetInterfaceIndex() != ifaceIndex)
+ continue;
+ ProtoAddress dst;
+ membership->GetDstAddr(dst);
+ if (dst.HostIsEqual(group))
+ return membership;
+ }
+ return NULL;
+}
+
+static const char* OifWhy(MulticastFIB::Membership* membership,
+ Smf::Interface* iface,
+ const ProtoAddress& group)
{
- if (useJson) mcast_fib.DumpFlowListJson(brief, ss);
- else mcast_fib.DumpFlowList(brief, ss);
+ if (NULL != membership)
+ {
+ if (membership->FlagIsSet(MulticastFIB::Membership::MANAGED))
+ return "managed";
+ if (membership->FlagIsSet(MulticastFIB::Membership::STATIC))
+ return "static";
+ if (membership->FlagIsSet(MulticastFIB::Membership::ELASTIC))
+ return "ack";
+ }
+ if ((NULL != iface) && iface->IsManaged() && iface->HasActiveMembership(group))
+ return "managed";
+ return "-";
+}
+
+static const char* OifMode(Smf::Interface* iface)
+{
+ if ((NULL != iface) && iface->GetElasticMulticast())
+ return "elastic";
+ return "cf";
+}
+
+static void DumpGroupsNested(MulticastFIB& fib, Smf& smf,
+ MulticastFIB::MembershipTable* memberships,
+ bool useJson, std::ostringstream& ss,
+ unsigned int currentTick)
+{
+ MulticastFIB::EntryTable::Iterator fiberator(fib.AccessFlowTable());
+ MulticastFIB::Entry* entry;
+ bool comma = false;
+ if (useJson)
+ ss << "[";
+ else
+ {
+ ss << "Flags: A = Active, K = Ack\n"
+ << "Selected: Y = current upstream, N = additional inbound\n"
+ << "Fwd: B = Block, H = Hybrid, L = Limit, F = Forward, D = Deny\n"
+ << "Why: managed = IGMP/static last hop, ack = EM_ACK, - = default\n\n"
+ << "group saddr flags ack recv pps\n"
+ << "---------------- ---------------- ----- --- ------ ----\n";
+ }
+ while (NULL != (entry = fiberator.GetNextEntry()))
+ {
+ ProtoAddress dst, src;
+ char group[64] = "-";
+ char saddr[64] = "*";
+ char flags[8];
+ entry->GetDstAddr(dst);
+ dst.GetHostString(group, sizeof(group) - 1);
+ entry->GetSrcAddr(src);
+ if (src.IsValid())
+ src.GetHostString(saddr, sizeof(saddr) - 1);
+ FormatFlowFlagLetters(entry->IsActive(), entry->GetAckingStatus(), flags, sizeof(flags));
+ MulticastFIB::UpstreamRelay* best = entry->GetCurrentBestUpstreamRelay();
+ char bestIif[Smf::IF_NAME_MAX + 1];
+ char bestUp[64] = "";
+ FormatIfaceName(best ? best->GetInterfaceIndex() : 0, bestIif, Smf::IF_NAME_MAX, "-");
+ if (best)
+ best->GetAddress().GetHostString(bestUp, sizeof(bestUp) - 1);
+
+ if (useJson)
+ {
+ ss << (comma ? "," : "") << "{";
+ ss << "\"MCastAddr\" : \"" << group << "\",";
+ ss << "\"SrcAddr\" : \"" << saddr << "\",";
+ ss << "\"Status\" : \"" << (entry->IsActive() ? "ACTIVE" : "IDLE") << "\",";
+ ss << "\"FwdStatus\" : \""
+ << MulticastFIB::GetForwardingStatusString(entry->GetDefaultForwardingStatus())
+ << "\",";
+ ss << "\"Ack\" : \"" << (entry->GetAckingStatus() ? "yes" : "no") << "\",";
+ ss << "\"RecvPkts\" : " << entry->GetRecvPkts() << ",";
+ ss << "\"RecvPps\" : " << std::fixed << std::setprecision(1)
+ << entry->GetRecvPps(currentTick) << ",";
+ ss << "\"SrcInterface\" : \"" << bestIif << "\",";
+ ss << "\"UpstreamAddr\" : \"" << bestUp << "\",";
+ ss << "\"Sources\" : [";
+ MulticastFIB::UpstreamRelayList::Iterator uperator(entry->AccessUpstreamRelayList());
+ MulticastFIB::UpstreamRelay* relay;
+ bool ucomma = false;
+ while (NULL != (relay = uperator.GetNextItem()))
+ {
+ char iif[Smf::IF_NAME_MAX + 1];
+ char uaddr[64] = "";
+ FormatIfaceName(relay->GetInterfaceIndex(), iif, Smf::IF_NAME_MAX);
+ relay->GetAddress().GetHostString(uaddr, sizeof(uaddr) - 1);
+ ss << (ucomma ? "," : "")
+ << "{\"Interface\" : \"" << iif
+ << "\", \"UpstreamAddr\" : \"" << uaddr
+ << "\", \"Selected\" : " << ((relay == best) ? "true" : "false")
+ << "}";
+ ucomma = true;
+ }
+ ss << "], \"Outbound\" : [";
+ MulticastFIB::BucketList::Iterator biterator(entry->AccessBucketList());
+ MulticastFIB::TokenBucket* bucket;
+ bool bcomma = false;
+ while (NULL != (bucket = biterator.GetNextItem()))
+ {
+ char oif[Smf::IF_NAME_MAX + 1];
+ FormatIfaceName(bucket->GetInterfaceIndex(), oif, Smf::IF_NAME_MAX);
+ Smf::Interface* iface = smf.GetInterface(bucket->GetInterfaceIndex());
+ MulticastFIB::Membership* mem = (NULL != memberships) ?
+ FindIfaceMembership(*memberships, bucket->GetInterfaceIndex(), dst) : NULL;
+ ss << (bcomma ? "," : "")
+ << "{\"Interface\" : \"" << oif
+ << "\", \"FwdStatus\" : \""
+ << MulticastFIB::GetForwardingStatusString(bucket->GetForwardingStatus())
+ << "\", \"Mode\" : \"" << OifMode(iface)
+ << "\", \"Why\" : \"" << OifWhy(mem, iface, dst)
+ << "\", \"SentPkts\" : " << bucket->GetSentPkts()
+ << ", \"SentPps\" : " << std::fixed << std::setprecision(1)
+ << bucket->GetSentPps(currentTick)
+ << ", \"DownstreamRelays\" : ";
+ if (NULL != mem)
+ DumpDownstreamRelays(*mem, ss, true);
+ else
+ ss << "[]";
+ ss << "}";
+ bcomma = true;
+ }
+ ss << "]}";
+ }
+ else
+ {
+ ss << std::left << std::setw(16) << group << " "
+ << std::setw(16) << saddr << " "
+ << std::setw(5) << flags << " "
+ << std::setw(3) << (entry->GetAckingStatus() ? "yes" : "no") << " "
+ << std::setw(6) << entry->GetRecvPkts() << " "
+ << std::fixed << std::setprecision(1) << std::setw(4)
+ << entry->GetRecvPps(currentTick) << "\n";
+ ss << " sources\n"
+ << " iif upstream sel\n";
+ MulticastFIB::UpstreamRelayList::Iterator uperator(entry->AccessUpstreamRelayList());
+ MulticastFIB::UpstreamRelay* relay;
+ bool anyUp = false;
+ while (NULL != (relay = uperator.GetNextItem()))
+ {
+ char iif[Smf::IF_NAME_MAX + 1];
+ char uaddr[64] = "-";
+ FormatIfaceName(relay->GetInterfaceIndex(), iif, Smf::IF_NAME_MAX);
+ relay->GetAddress().GetHostString(uaddr, sizeof(uaddr) - 1);
+ ss << " " << std::left << std::setw(12) << iif << " "
+ << std::setw(16) << uaddr << " "
+ << ((relay == best) ? "Y" : "N") << "\n";
+ anyUp = true;
+ }
+ if (!anyUp)
+ ss << " -\n";
+ ss << " downstream\n"
+ << " oif mode why fwd sent pps relays\n";
+ MulticastFIB::BucketList::Iterator biterator(entry->AccessBucketList());
+ MulticastFIB::TokenBucket* bucket;
+ bool anyOif = false;
+ while (NULL != (bucket = biterator.GetNextItem()))
+ {
+ char oif[Smf::IF_NAME_MAX + 1];
+ FormatIfaceName(bucket->GetInterfaceIndex(), oif, Smf::IF_NAME_MAX);
+ Smf::Interface* iface = smf.GetInterface(bucket->GetInterfaceIndex());
+ MulticastFIB::Membership* mem = (NULL != memberships) ?
+ FindIfaceMembership(*memberships, bucket->GetInterfaceIndex(), dst) : NULL;
+ ss << " " << std::left << std::setw(12) << oif << " "
+ << std::setw(8) << OifMode(iface) << " "
+ << std::setw(8) << OifWhy(mem, iface, dst) << " "
+ << FwdStatusLetter(bucket->GetForwardingStatus()) << " "
+ << std::setw(6) << bucket->GetSentPkts() << " "
+ << std::fixed << std::setprecision(1) << std::setw(4)
+ << bucket->GetSentPps(currentTick) << " ";
+ if (NULL != mem)
+ DumpDownstreamRelays(*mem, ss, false);
+ else
+ ss << "-";
+ ss << "\n";
+ anyOif = true;
+ }
+ if (!anyOif)
+ ss << " -\n";
+ }
+ comma = true;
+ }
+ if (useJson)
+ ss << "]\n";
+}
+
+void ElasticMulticastForwarder::DumpGroups(bool brief, bool useJson, std::ostringstream& ss, bool details)
+{
+ unsigned int now = UpdateTicker();
+ if (details && !brief)
+ {
+ Smf& smf = *reinterpret_cast(this);
+ MulticastFIB::MembershipTable* memberships =
+ (NULL != mcast_controller) ? &mcast_controller->AccessMembershipTable() : NULL;
+ DumpGroupsNested(mcast_fib, smf, memberships, useJson, ss, now);
+ return;
+ }
+ if (useJson) mcast_fib.DumpFlowListJson(brief, ss, details, now);
+ else mcast_fib.DumpFlowList(brief, ss, details, now);
+}
+
+void ElasticMulticastForwarder::DumpManagedGroups(bool useJson, std::ostringstream& ss)
+{
+ Smf& smf = *reinterpret_cast(this);
+ Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
+ Smf::Interface* iface;
+ bool comma = false;
+ if (useJson)
+ ss << "[";
+ else
+ {
+ ss << "Flags: M = Managed (last-hop HasActiveMembership)\n"
+ << "iface flags groups\n"
+ << "------------ ----- ------\n";
+ }
+ while (NULL != (iface = iterator.GetNextItem()))
+ {
+ if (useJson)
+ {
+ ss << (comma ? "," : "")
+ << "{\"Interface\" : \"" << iface->GetNameStr()
+ << "\", \"Managed\" : " << (iface->IsManaged() ? "true" : "false")
+ << ", \"Groups\" : [";
+ ProtoAddressList::Iterator git(iface->AccessManagedMemberships());
+ ProtoAddress grp;
+ bool gcomma = false;
+ while (git.GetNextAddress(grp))
+ {
+ char host[64];
+ grp.GetHostString(host, sizeof(host) - 1);
+ ss << (gcomma ? "," : "") << "\"" << host << "\"";
+ gcomma = true;
+ }
+ ss << "]}";
+ }
+ else
+ {
+ ss << std::left << std::setw(14) << iface->GetNameStr()
+ << (iface->IsManaged() ? "M " : "- ");
+ ProtoAddressList::Iterator git(iface->AccessManagedMemberships());
+ ProtoAddress grp;
+ bool gcomma = false;
+ while (git.GetNextAddress(grp))
+ {
+ char host[64];
+ grp.GetHostString(host, sizeof(host) - 1);
+ ss << (gcomma ? "," : "") << host;
+ gcomma = true;
+ }
+ if (!gcomma)
+ ss << "-";
+ ss << "\n";
+ }
+ comma = true;
+ }
+ if (useJson)
+ ss << "]\n";
}
///////////////////////////////////////////////////////////////////////////////////
@@ -2493,11 +2948,21 @@ bool ElasticMulticastController::AddManagedMembership(const ProtoFlow::Descripti
membership->GetFlowDescription().Print();
PLOG(PL_ALWAYS, "\n");
}
- if (0 == membership->GetFlags())
- {
- mcast_forwarder->SetAckingStatus(membership->GetFlowDescription(), true);
- mcast_forwarder->SetForwardingStatus(membership->GetFlowDescription(), ifaceIndex, MulticastFIB::FORWARD, true);
- }
+ // Always push FORWARD/ack. A first call can race the FIB entry
+ // (SetForwardingStatus is a no-op if the flow is not there yet);
+ // skipping later refreshes when flags are already set leaves the
+ // last hop stuck at LIMIT.
+ // Flow key must omit iface index: packet FIB entries have index 0.
+ // Membership includes the host iface, so FLAG_ALL on that
+ // description matches nothing and eth1 stays LIMIT.
+ ProtoAddress dstIp, srcIp;
+ membership->GetFlowDescription().GetDstAddr(dstIp);
+ membership->GetFlowDescription().GetSrcAddr(srcIp);
+ ProtoFlow::Description fwdDesc(dstIp, srcIp,
+ membership->GetFlowDescription().GetTrafficClass(),
+ membership->GetFlowDescription().GetProtocol());
+ mcast_forwarder->SetAckingStatus(fwdDesc, true);
+ mcast_forwarder->SetForwardingStatus(fwdDesc, ifaceIndex, MulticastFIB::FORWARD, true);
// Set MANAGED status for _all_ matching memberships for this "ifaceIndex"
MulticastFIB::MembershipTable::Iterator iterator(membership_table, &membership->GetFlowDescription());
while (NULL != (membership = iterator.GetNextEntry()))
@@ -2722,7 +3187,8 @@ void ElasticMulticastController::HandleAck(const ElasticAck& ack,
{
//auto smf=reinterpret_cast(mcast_forwarder);
Smf::Interface* iface = static_cast(smfIface);
- if (iface->GetIpAddress().GetType() == ProtoAddress::INVALID)
+ if ((iface->GetIpAddress().GetType() == ProtoAddress::INVALID) &&
+ !iface->GetTunnelLocalAddress().IsValid())
{
PLOG(PL_WARN, "ElasticMulticastController::HandleAck() no IP address on interface %s!\n", iface->GetNameStr());
return;
@@ -3030,15 +3496,22 @@ void ElasticMulticastController::Update(const ProtoFlow::Description& flowDescr
if (ignoreIdleCount) return; // NOT SURE THIS ACTUALLY WORKS if ignoreIdleCount == true
// Iterate across all matching (per-interface) memberships, update the packet
- // counts and status for "ELASTIC" memberships as appropriate
- // NOTE: this iterator finds _all_ matching interfaces, including dst-only
- MulticastFIB::MembershipTable::Iterator iterator(membership_table, &flowDescription);
+ // counts and status for "ELASTIC" memberships as appropriate.
+ // FLAG_DST so dst-only IGMP (*,G) memberships match a packet (S,G).
+ // FLAG_ALL prefixes the source and never visits those entries.
+ MulticastFIB::MembershipTable::Iterator iterator(membership_table, &flowDescription,
+ ProtoFlow::Description::FLAG_DST);
bool ackingStatus = false;
MulticastFIB::Membership* membership;
while (NULL != (membership = iterator.GetNextEntry()))
{
if (membership->FlagIsSet(MulticastFIB::Membership::ELASTIC))
{
+ ProtoAddress memSrc, pktSrc;
+ membership->GetFlowDescription().GetSrcAddr(memSrc);
+ flowDescription.GetSrcAddr(pktSrc);
+ if (memSrc.IsValid() && pktSrc.IsValid() && !memSrc.HostIsEqual(pktSrc))
+ continue;
unsigned int totalPktCount = membership->IncrementIdleCount(pktCount);
// set the idle threshold to the pps for the flow
// updateInterval is microseconds
@@ -3103,6 +3576,17 @@ void ElasticMulticastController::Update(const ProtoFlow::Description& flowDescr
else
{
ackingStatus = true;
+ // MANAGED/STATIC (IGMP): the packet flow is the iterator key
+ // (S,G). Push FORWARD on that host iface here; AddManagedMembership
+ // can miss if it searches with a (*,G) membership description.
+ if (membership->FlagIsSet(MulticastFIB::Membership::MANAGED) ||
+ membership->FlagIsSet(MulticastFIB::Membership::STATIC))
+ {
+ mcast_forwarder->SetForwardingStatus(flowDescription,
+ membership->GetInterfaceIndex(),
+ MulticastFIB::FORWARD,
+ true);
+ }
}
}
if (ackingStatus != oldAckingStatus)
@@ -3119,8 +3603,73 @@ void ElasticMulticastController::Update(const ProtoFlow::Description& flowDescr
- void ElasticMulticastController::DumpGroups(bool brief, bool useJson, std::ostringstream& ss)
+void ElasticMulticastController::DumpGroups(bool brief, bool useJson, std::ostringstream& ss, bool details)
{
- mcast_forwarder->DumpGroups(brief, useJson, ss);
+ mcast_forwarder->DumpGroups(brief, useJson, ss, details);
+}
+
+void ElasticMulticastController::DumpManagedGroups(bool useJson, std::ostringstream& ss)
+{
+ mcast_forwarder->DumpManagedGroups(useJson, ss);
+}
+
+void ElasticMulticastController::DumpMemberships(bool useJson, std::ostringstream& ss)
+{
+ MulticastFIB::MembershipTable::Iterator iterator(membership_table);
+ MulticastFIB::Membership* membership;
+ bool comma = false;
+
+ if (useJson)
+ ss << "[";
+ else
+ {
+ ss << "Flags: S = Static, M = Managed, E = Elastic\n"
+ << "group saddr iface flags relays\n"
+ << "---------------- ---------------- ------------ ----- ----------------\n";
+ }
+
+ while (NULL != (membership = iterator.GetNextEntry()))
+ {
+ ProtoAddress dst, src;
+ char dstHost[256] = "";
+ char srcHost[256] = "*";
+ char ifaceName[Smf::IF_NAME_MAX + 1];
+
+ membership->GetDstAddr(dst);
+ dst.GetHostString(dstHost, 255);
+ membership->GetSrcAddr(src);
+ if (src.IsValid())
+ src.GetHostString(srcHost, 255);
+ FormatIfaceName(membership->GetInterfaceIndex(), ifaceName, Smf::IF_NAME_MAX);
+
+ if (useJson)
+ {
+ ss << (comma ? "," : "") << "{";
+ ss << "\"MCastAddr\" : \"" << dstHost << "\",";
+ ss << "\"SrcAddr\" : \"" << srcHost << "\",";
+ ss << "\"Interface\" : \"" << ifaceName << "\",";
+ ss << "\"Flags\" : ";
+ DumpMembershipFlags(membership->GetFlags(), ss);
+ ss << ", \"DownstreamRelays\" : ";
+ DumpDownstreamRelays(*membership, ss, true);
+ ss << "}";
+ comma = true;
+ }
+ else
+ {
+ char flags[8];
+ FormatMembershipFlagLetters(membership->GetFlags(), flags, sizeof(flags));
+ FormatIfaceName(membership->GetInterfaceIndex(), ifaceName, Smf::IF_NAME_MAX, "-");
+ ss << std::left << std::setw(16) << dstHost << " ";
+ ss << std::setw(16) << srcHost << " ";
+ ss << std::setw(12) << ifaceName << " ";
+ ss << std::setw(5) << flags << " ";
+ DumpDownstreamRelays(*membership, ss, false);
+ ss << "\n";
+ }
+ }
+
+ if (useJson)
+ ss << "]\n";
}
diff --git a/src/common/nrlsmf.cpp b/src/common/nrlsmf.cpp
index fbf91f9..4ce8ecd 100644
--- a/src/common/nrlsmf.cpp
+++ b/src/common/nrlsmf.cpp
@@ -53,6 +53,9 @@
#include
#include
#include
+#ifdef LINUX
+#include
+#endif
#include
#include
#include
@@ -81,6 +84,8 @@ class SmfApp : public ProtoApp
// This is used by ElasticMulticastForwarder to send EM_ACKs, etc
bool SendFrame(unsigned int ifaceIndex, char* buffer, unsigned int length);
+ bool SendFrameTo(unsigned int ifaceIndex, char* buffer, unsigned int length,
+ const ProtoAddress& dest);
private:
void MonitorEventHandler(ProtoChannel& theChannel,
@@ -100,6 +105,21 @@ class SmfApp : public ProtoApp
static void CliUsage();
static int RunControlClient(int argc, const char*const* argv);
+ bool ControlReply(const char* data, unsigned int numBytes);
+ bool ControlReply(const std::string& s);
+ void OnShowCommand(const char* arg);
+ void ReplyVersion(bool json);
+ void ReplyStats(bool json);
+ void ReplyInfo(bool json);
+ void ReplyInterfaces(bool json);
+ void ReplyTunnel(bool json);
+ void ReplyTunnelNeighbors(bool json);
+#ifdef ELASTIC_MCAST
+ void ReplyGroups(bool json, bool brief, bool details = false);
+ void ReplyGroupMemberships(bool json);
+ void ReplyIgmpGroups(bool json);
+#endif // ELASTIC_MCAST
+
bool LoadConfig(const char* configPath);
bool ProcessGroupConfig(ProtoJson::Object& groupConfig);
bool ProcessInterfaceConfig(ProtoJson::Object& ifaceConfig);
@@ -195,6 +215,24 @@ class SmfApp : public ProtoApp
bool JoinUnderlayGroup(const ProtoAddress& groupAddr, const char* ifaceName);
bool LeaveUnderlayGroup(const ProtoAddress& groupAddr, const char* ifaceName);
+ bool GreDeviceIsUnicastMgre(Smf::Interface& iface);
+ bool MaybeEnableUnderlayGreDemux(const ProtoAddress& groupAddr);
+ bool EnableUnderlayGreDemux(const ProtoAddress& groupAddr, const char* ifaceName);
+ void OnUnderlayGreCapture(ProtoChannel& theChannel,
+ ProtoChannel::Notification notifyType);
+
+ static bool OnNeighborDump(unsigned int ifIndex,
+ const ProtoAddress& dst,
+ const ProtoAddress& lladdr,
+ unsigned short ndmState,
+ void* userData);
+ void DumpLearnedTunnelNeighbors(Smf::Interface& iface);
+ void ClearLearnedTunnelNeighbors(Smf::Interface& iface);
+ void UpdateLearnedTunnelPeer(Smf::Interface& iface,
+ const ProtoAddress& overlay,
+ const ProtoAddress& underlay,
+ unsigned short ndmState,
+ bool deleted);
void DisplayGroups();
@@ -240,7 +278,7 @@ class SmfApp : public ProtoApp
class InterfaceMechanism : public Smf::Interface::Extension
{
public:
- InterfaceMechanism(Smf::Interface& iface, SmfPacket::Pool& pktPool);
+ InterfaceMechanism(Smf::Interface& iface, SmfPacket::Pool& pktPool, Smf& theSmf);
~InterfaceMechanism();
Smf::Interface& GetInterface() {return smf_iface;}
@@ -283,6 +321,9 @@ class SmfApp : public ProtoApp
enum TxStatus {TX_OK, TX_BLOCK,TX_ERROR};
TxStatus SendFrame(char* frame, unsigned int frameLen);
+ bool SendGrePayload(ProtoCap& cap, char* frame, unsigned int frameLength, unsigned int& numBytes);
+ bool SendGreToRemote(ProtoCap& cap, char* frame, unsigned int frameLength,
+ const ProtoAddress& dest, unsigned int& numBytes);
void ResetTxIterator() {tx_iterator.Reset();}
@@ -307,6 +348,7 @@ class SmfApp : public ProtoApp
private:
Smf::Interface& smf_iface;
SmfPacket::Pool& pkt_pool;
+ Smf& smf;
ProtoVif* proto_vif;
bool is_shadowing;
bool block_igmp;
@@ -426,6 +468,9 @@ class SmfApp : public ProtoApp
ProtoTimer igmp_query_timer;
#endif // ELASTIC_MCAST
ProtoSocket underlay_group_socket; // used for joining mGRE underlay groups
+ ProtoAddressList underlay_join_groups; // ujoin group -> underlay ifindex
+ ProtoCap* underlay_gre_cap; // GRE-in-mcast capture (unicast mGRE only)
+ ProtoAddressList underlay_gre_groups; // groups that need capture demux
#ifdef ADAPTIVE_ROUTING
SmartController smart_controller;
#endif // ADAPTIVE_ROUTING
@@ -455,8 +500,8 @@ class SmfApp : public ProtoApp
const unsigned int SmfApp::BUFFER_MAX = FRAME_SIZE_MAX + 2 + (256 *sizeof(UINT32));
-SmfApp::InterfaceMechanism::InterfaceMechanism(Smf::Interface& iface, SmfPacket::Pool& pktPool)
- : smf_iface(iface), pkt_pool(pktPool), proto_vif(NULL), is_shadowing(false), block_igmp(false),
+SmfApp::InterfaceMechanism::InterfaceMechanism(Smf::Interface& iface, SmfPacket::Pool& pktPool, Smf& theSmf)
+ : smf_iface(iface), pkt_pool(pktPool), smf(theSmf), proto_vif(NULL), is_shadowing(false), block_igmp(false),
cid_list_length(0), cid_mirror(true), tx_iterator(cid_list), output_notification(false),
#ifdef _PROTO_DETOUR
proto_detour(NULL),
@@ -650,6 +695,65 @@ SmfApp::CidElement* SmfApp::InterfaceMechanism::GetNextTxElement(bool autoReset)
return elem;
} // end SmfApp::InterfaceMechanism::GetNetTxElement()
+bool SmfApp::InterfaceMechanism::SendGrePayload(ProtoCap& cap, char* frame, unsigned int frameLength, unsigned int& numBytes)
+{
+ // GRE inject is inner IP only. A wildcard-remote mGRE interface has no
+ // single kernel dest; overlay multicast is sent once per mapped remote
+ // (unicast peers and/or an underlay multicast group).
+ numBytes = frameLength - 14;
+ char* payload = frame + 14;
+ const bool overlayMcast = (0 != (0x01 & ((UINT8)frame[0])));
+ ProtoAddressList dests;
+ if (overlayMcast)
+ smf.GetTunnelUnicastRemotes(smf_iface.GetIndex(), dests);
+ if (dests.IsEmpty())
+ return cap.Send(payload, numBytes);
+
+ ProtoAddress saved = cap.GetTunnelRemoteAddr();
+ bool success = false;
+ unsigned int sentBytes = 0;
+ ProtoAddressList::Iterator it(dests);
+ ProtoAddress peer;
+ while (it.GetNextAddress(peer))
+ {
+ unsigned int n = frameLength - 14;
+ cap.SetTunnelRemoteAddr(peer);
+ bool ok = cap.Send(payload, n);
+ if (0 == n)
+ ok = false;
+ if (ok)
+ {
+ success = true;
+ sentBytes = n;
+ }
+ else
+ {
+ PLOG(PL_WARN, "SendGrePayload() failed via %s dest %s\n",
+ smf_iface.GetNameStr(), peer.GetHostString());
+ }
+ }
+ cap.SetTunnelRemoteAddr(saved);
+ numBytes = success ? sentBytes : 0;
+ return success;
+} // end SmfApp::InterfaceMechanism::SendGrePayload()
+
+bool SmfApp::InterfaceMechanism::SendGreToRemote(ProtoCap& cap, char* frame, unsigned int frameLength,
+ const ProtoAddress& dest, unsigned int& numBytes)
+{
+ // One GRE outer dest (EM_ACK to a specific mGRE neighbor).
+ numBytes = frameLength - 14;
+ char* payload = frame + 14;
+ ProtoAddress saved = cap.GetTunnelRemoteAddr();
+ cap.SetTunnelRemoteAddr(dest);
+ unsigned int n = numBytes;
+ bool ok = cap.Send(payload, n);
+ cap.SetTunnelRemoteAddr(saved);
+ if (0 == n)
+ ok = false;
+ numBytes = ok ? n : 0;
+ return ok;
+} // end SmfApp::InterfaceMechanism::SendGreToRemote()
+
SmfApp::InterfaceMechanism::TxStatus SmfApp::InterfaceMechanism::SendFrame(char* frame, unsigned int frameLength)
{
bool success = false;
@@ -665,9 +769,7 @@ SmfApp::InterfaceMechanism::TxStatus SmfApp::InterfaceMechanism::SendFrame(char*
}
else if (ProtoNet::IFACE_GRE == elem->GetProtoCap().GetInterfaceType())
{
- // Just send the IP payload portion
- numBytes -= 14;
- success = elem->GetProtoCap().Send(frame + 14, numBytes);
+ success = SendGrePayload(elem->GetProtoCap(), frame, frameLength, numBytes);
}
else if ((NULL != proto_vif) && !is_shadowing)
{
@@ -697,8 +799,7 @@ SmfApp::InterfaceMechanism::TxStatus SmfApp::InterfaceMechanism::SendFrame(char*
bool mirrorSuccess;
if (ProtoNet::IFACE_GRE == elem->GetProtoCap().GetInterfaceType())
{
- numBytes -= 14;
- mirrorSuccess = elem->GetProtoCap().Send(frame + 14, numBytes);
+ mirrorSuccess = SendGrePayload(elem->GetProtoCap(), frame, frameLength, numBytes);
}
else if (is_shadowing)
{
@@ -730,8 +831,7 @@ SmfApp::InterfaceMechanism::TxStatus SmfApp::InterfaceMechanism::SendFrame(char*
numBytes = frameLength;
if (ProtoNet::IFACE_GRE == elem->GetProtoCap().GetInterfaceType())
{
- numBytes -= 14;
- success = elem->GetProtoCap().Send(frame + 14, numBytes);
+ success = SendGrePayload(elem->GetProtoCap(), frame, frameLength, numBytes);
}
else if (is_shadowing)
{
@@ -997,6 +1097,7 @@ SmfApp::SmfApp()
igmp_controller(GetTimerMgr(), smf),
#endif // ELASTIC_MCAST
underlay_group_socket(ProtoSocket::UDP),
+ underlay_gre_cap(NULL),
#ifdef ADAPTIVE_ROUTING
smart_controller(GetTimerMgr()),
#endif // ADAPTIVE_ROUTING
@@ -1039,7 +1140,7 @@ void SmfApp::Usage()
{
const char* const* nextCmd = CMD_LIST;
fprintf(stderr, "Usage: nrlsmf [options]:\n");
- fprintf(stderr, " nrlsmf --cli [-i ] [args...]\n");
+ fprintf(stderr, " nrlsmf --cli [-i ] -c [-c ...]\n");
while (*nextCmd) {
const char* cmd = &(*nextCmd)[1];
nextCmd++;
@@ -1112,7 +1213,7 @@ const char* const SmfApp::CMD_LIST[] =
"+leave", "[->][,[,]]] (Note can optionally be an interface name)",
"+load", " : load nrlsmf JSON configuration file",
"+log", " : debug log file",
- "+map", ",[,] : maps tunnel endpoint information for a GRE interface (or other ancillary iface->address assocation)",
+ "+map", ",[,|dynamic] : map tunnel endpoints (0.0.0.0 = wildcard remote; dynamic = learn peers from kernel neigh)",
"+merge", " : forward _among_ all iface's listed",
"+push", " : forward packets from srcIFace to all dstIface's listed",
"+queue", "[,] : perform SMF packet queuing",
@@ -1176,75 +1277,197 @@ SmfApp::CmdType SmfApp::GetCmdType(const char* cmd)
return type;
} // end SmfApp::GetCmdType()
-// Status/query verbs handled in OnControlMsg that reply on server_pipe.
-static const struct
+// Modern CLI show commands. Legacy one-shot verbs (groupsj, jsonStats, ...)
+// remain on the control socket for machine clients.
+// A command is plus an optional subcommand (e.g. "interface grouping").
+// json/brief/details are optional modifiers; a command may support none, one, or
+// both of brief/details. The unmodified command is the default listing.
+static const struct ShowTopicSpec
{
const char* name;
+ const char* sub; // NULL, or a subcommand such as "grouping"
const char* help;
+ bool json;
+ bool brief;
+ bool details;
bool elasticOnly;
-} kCliQueryCmds[] =
+} kShowTopics[] =
{
- { "ping", "heartbeat; returns pong if nrlsmf is running", false },
- { "stats", "per-interface packet/flow counters (text table)", false },
- { "jsonStats", "same as stats, JSON", false },
- { "info", "interface groups (text table)", false },
- { "jsonInfo", "same as info, JSON", false },
- { "jsonVersion", "nrlsmf version as JSON", false },
- { "interfaces", "configured interfaces (text table)", false },
- { "interfacesj", "same as interfaces, JSON", false },
- { "groups", "elastic multicast groups (text)", true },
- { "groupsj", "same as groups, JSON", true },
- { "brfgroups", "brief elastic multicast groups (text)", true },
- { "brfgroupsj", "same as brfgroups, JSON", true },
- { NULL, NULL, false }
+ { "version", NULL, "nrlsmf version", true, false, false, false },
+ { "statistics", NULL, "per-interface packet/flow counters", true, false, false, false },
+ { "interface", NULL, "configured interfaces", true, false, false, false },
+ { "interface", "grouping", "SMF interface groups", true, false, false, false },
+ { "tunnel", NULL, "tunnel endpoint mappings", true, false, false, false },
+ { "tunnel", "neighbors", "GRE/mGRE neighbors", true, false, false, false },
+ { "groups", NULL, "elastic multicast flows", true, true, true, true },
+ { "groups", "memberships", "EM memberships and downstream relays", true, false, false, true },
+ { "igmp", "groups", "IGMP last-hop groups (with-frr)", true, false, false, true },
+ { NULL, NULL, NULL, false, false, false, false }
};
-static void CliQueryHelp()
+static const ShowTopicSpec* FindShowTopic(const char* name, const char* sub)
{
- printf("Query / show commands:\n");
- printf(" nrlsmf --cli [-i ] \n\n");
- for (unsigned int i = 0; NULL != kCliQueryCmds[i].name; i++)
+ if (NULL == name)
+ return NULL;
+ for (unsigned int i = 0; NULL != kShowTopics[i].name; i++)
{
- printf(" %-14s %s%s\n",
- kCliQueryCmds[i].name,
- kCliQueryCmds[i].help,
- kCliQueryCmds[i].elasticOnly ? " (elastic build)" : "");
+ if (0 != strcmp(name, kShowTopics[i].name))
+ continue;
+ if (NULL == sub)
+ {
+ if (NULL == kShowTopics[i].sub)
+ return &kShowTopics[i];
+ }
+ else if ((NULL != kShowTopics[i].sub) && (0 == strcmp(sub, kShowTopics[i].sub)))
+ {
+ return &kShowTopics[i];
+ }
}
- printf("\nConfiguration commands (debug, add, relay, map, ...) use the same\n"
- "--cli syntax and do not return a reply. See \"nrlsmf help\".\n");
+ return NULL;
+}
+
+static bool IsShowSubcommand(const char* name, const char* word)
+{
+ return (NULL != FindShowTopic(name, word));
+}
+
+static void FormatShowMods(std::ostringstream& ss, const ShowTopicSpec& topic)
+{
+ if (topic.brief)
+ ss << " [brief]";
+ if (topic.details)
+ ss << " [details]";
+ if (topic.json)
+ ss << " [json]";
+}
+
+static void FormatShowHelp(std::ostringstream& ss)
+{
+ ss << "Show commands:\n"
+ << " nrlsmf --cli [-i ] -c \"show [modifiers]\"\n"
+ << "\n";
+ for (unsigned int i = 0; NULL != kShowTopics[i].name; i++)
+ {
+ std::ostringstream line;
+ line << " show " << kShowTopics[i].name;
+ if (NULL != kShowTopics[i].sub)
+ line << " " << kShowTopics[i].sub;
+ FormatShowMods(line, kShowTopics[i]);
+ ss << std::left << std::setw(40) << line.str() << kShowTopics[i].help;
+ if (kShowTopics[i].elasticOnly)
+ ss << " (elastic build)";
+ ss << "\n";
+ }
+ ss << "\n"
+ << "Modifiers are optional and command-specific. json, when used, is last:\n"
+ << " brief less output than the default listing\n"
+ << " details more output than the default listing\n"
+ << " json machine-readable JSON\n"
+ << "\n"
+ << "Examples:\n"
+ << " nrlsmf --cli -c \"show statistics\"\n"
+ << " nrlsmf --cli -c \"show statistics json\"\n"
+ << " nrlsmf --cli -c \"show interface\"\n"
+ << " nrlsmf --cli -c \"show interface grouping\"\n"
+ << " nrlsmf --cli -c \"show interface grouping json\"\n"
+ << " nrlsmf --cli -c \"show tunnel\"\n"
+ << " nrlsmf --cli -c \"show tunnel neighbors\"\n"
+ << " nrlsmf --cli -c \"show groups brief\"\n"
+ << " nrlsmf --cli -c \"show groups brief json\"\n"
+ << " nrlsmf --cli -c \"show groups details json\"\n"
+ << " nrlsmf --cli -c \"show groups memberships json\"\n"
+ << " nrlsmf --cli -c \"show igmp groups json\"\n"
+ << " nrlsmf --cli -i smf-p4 -c \"show interface json\"\n"
+ << "\n"
+ << "Configuration commands (debug, add, relay, map, ...) use the same\n"
+ << "--cli -c syntax and do not return a reply. See \"nrlsmf help\".\n";
+}
+
+static void CliQueryHelp()
+{
+ std::ostringstream ss;
+ FormatShowHelp(ss);
+ fputs(ss.str().c_str(), stdout);
}
void SmfApp::CliUsage()
{
fprintf(stderr,
- "Usage: nrlsmf --cli [-i ] [args...]\n"
+ "Usage: nrlsmf --cli [-i ] -c [-c ...]\n"
" nrlsmf --cli [-i ] ?\n"
"\n"
"Send a runtime command to a running nrlsmf instance via its\n"
"control socket (default instance \"%s\" -> /tmp/%s).\n"
+ "Repeat -c to send more than one command in order.\n"
"\n"
"Options:\n"
" -i, --instance Target instance (default: %s)\n"
+ " -c, --cmd Command to send (repeatable)\n"
" -h, --help Show this help\n"
- " ? List query / show commands\n"
+ " ? List show commands\n"
"\n"
"Examples:\n"
" nrlsmf --cli ?\n"
- " nrlsmf --cli ping\n"
- " nrlsmf --cli stats\n"
- " nrlsmf --cli jsonInfo\n"
- " nrlsmf --cli debug 2\n"
- " nrlsmf --cli -i smf-r0 relay off\n",
+ " nrlsmf --cli -c \"show statistics\"\n"
+ " nrlsmf --cli -c \"show interface grouping\" -c \"show statistics\"\n"
+ " nrlsmf --cli -c \"show groups brief json\"\n"
+ " nrlsmf --cli -c \"debug 2\"\n"
+ " nrlsmf --cli -i smf-r0 -c \"relay off\"\n",
DEFAULT_INSTANCE_NAME, DEFAULT_INSTANCE_NAME, DEFAULT_INSTANCE_NAME);
}
-static bool CliExpectsReply(const char* cmd)
+static const char* CliSkipSpace(const char* s)
+{
+ while ((NULL != s) && isspace((unsigned char)*s))
+ s++;
+ return s;
+}
+
+static bool CliCopyFirstToken(const char* s, char* out, size_t outLen)
+{
+ s = CliSkipSpace(s);
+ if ((NULL == s) || ('\0' == *s) || (outLen < 2))
+ return false;
+ size_t n = 0;
+ while ((s[n] != '\0') && !isspace((unsigned char)s[n]))
+ n++;
+ if (n >= outLen)
+ n = outLen - 1;
+ memcpy(out, s, n);
+ out[n] = '\0';
+ return true;
+}
+
+static bool CliIsLocalHelp(const char* message)
+{
+ const char* s = CliSkipSpace(message);
+ if ((NULL == s) || ('\0' == *s))
+ return false;
+ if (0 == strcmp(s, "?"))
+ return true;
+ if (0 != strncmp(s, "show", 4) || ((s[4] != '\0') && !isspace((unsigned char)s[4])))
+ return false;
+ s = CliSkipSpace(s + 4);
+ return ('\0' == *s) || (0 == strcmp(s, "?"));
+}
+
+static bool CliExpectsReply(const char* message)
{
- if (NULL == cmd)
+ char cmd[64];
+ if (!CliCopyFirstToken(message, cmd, sizeof(cmd)))
return false;
- for (unsigned int i = 0; NULL != kCliQueryCmds[i].name; i++)
+ if ((0 == strcmp(cmd, "show")) || (0 == strcmp(cmd, "ping")))
+ return true;
+ static const char* const kLegacyQueryCmds[] =
+ {
+ "stats", "jsonStats", "info", "jsonInfo", "jsonVersion",
+ "interfaces", "interfacesj", "groups", "groupsj",
+ "brfgroups", "brfgroupsj",
+ NULL
+ };
+ for (const char* const* p = kLegacyQueryCmds; NULL != *p; p++)
{
- if (0 == strcmp(cmd, kCliQueryCmds[i].name))
+ if (0 == strcmp(cmd, *p))
return true;
}
return false;
@@ -1272,6 +1495,8 @@ int SmfApp::RunControlClient(int argc, const char*const* argv)
SetDebugLevel(PL_ERROR);
const char* instance = DEFAULT_INSTANCE_NAME;
+ std::vector commands;
+ bool sawHelpPositional = false;
int i = 2;
while (i < argc)
{
@@ -1291,6 +1516,28 @@ int SmfApp::RunControlClient(int argc, const char*const* argv)
instance = argv[++i];
i++;
}
+ else if ((0 == strcmp(argv[i], "-c")) || (0 == strcmp(argv[i], "--cmd")))
+ {
+ if ((i + 1) >= argc)
+ {
+ fprintf(stderr, "nrlsmf --cli: %s requires a command string\n", argv[i]);
+ CliUsage();
+ return 1;
+ }
+ const char* cmd = CliSkipSpace(argv[++i]);
+ if ((NULL == cmd) || ('\0' == *cmd))
+ {
+ fprintf(stderr, "nrlsmf --cli: empty command\n");
+ return 1;
+ }
+ commands.push_back(cmd);
+ i++;
+ }
+ else if (0 == strcmp(argv[i], "?"))
+ {
+ sawHelpPositional = true;
+ i++;
+ }
else if ('-' == argv[i][0])
{
fprintf(stderr, "nrlsmf --cli: unknown option %s\n", argv[i]);
@@ -1299,39 +1546,43 @@ int SmfApp::RunControlClient(int argc, const char*const* argv)
}
else
{
- break;
+ fprintf(stderr, "nrlsmf --cli: pass commands with -c, e.g. -c \"%s\"\n", argv[i]);
+ CliUsage();
+ return 1;
}
}
- if (i >= argc)
+ if (commands.empty())
{
+ if (sawHelpPositional)
+ {
+ CliQueryHelp();
+ return 0;
+ }
CliUsage();
return 1;
}
-
- if (0 == strcmp(argv[i], "?"))
+ if (sawHelpPositional)
{
- CliQueryHelp();
- return 0;
+ fprintf(stderr, "nrlsmf --cli: unexpected '?'; use -c \"?\" to list show commands\n");
+ return 1;
}
- const char* cmd = argv[i];
- char message[8192];
- size_t used = 0;
- message[0] = '\0';
- for (int a = i; a < argc; a++)
+ bool anyRemote = false;
+ for (size_t n = 0; n < commands.size(); n++)
{
- size_t argLen = strlen(argv[a]);
- if ((used + argLen + 2) >= sizeof(message))
+ if (!CliIsLocalHelp(commands[n]))
{
- fprintf(stderr, "nrlsmf --cli: command too long\n");
- return 1;
+ anyRemote = true;
+ break;
}
- if (used > 0)
- message[used++] = ' ';
- memcpy(message + used, argv[a], argLen);
- used += argLen;
- message[used] = '\0';
+ }
+
+ if (!anyRemote)
+ {
+ for (size_t n = 0; n < commands.size(); n++)
+ CliQueryHelp();
+ return 0;
}
char listenName[64];
@@ -1358,45 +1609,54 @@ int SmfApp::RunControlClient(int argc, const char*const* argv)
return 1;
}
- const bool wantsReply = CliExpectsReply(cmd);
- if (wantsReply)
+ bool startedServer = false;
+ for (size_t n = 0; n < commands.size(); n++)
{
- // Status replies are sent on server_pipe, not back to the requester.
- // Register this process as the (temporary) controller so we receive them.
- char startMsg[128];
- snprintf(startMsg, sizeof(startMsg), "smfServerStart %s", listenName);
- unsigned int numBytes = (unsigned int)strlen(startMsg) + 1;
- if (!smfPipe.Send(startMsg, numBytes))
+ const char* cmd = commands[n];
+ if (CliIsLocalHelp(cmd))
{
- fprintf(stderr, "nrlsmf --cli: failed to send smfServerStart to instance \"%s\"\n",
- instance);
- smfPipe.Close();
- listenPipe.Close();
- return 1;
+ CliQueryHelp();
+ continue;
}
- }
- unsigned int numBytes = (unsigned int)used + 1;
- if (!smfPipe.Send(message, numBytes))
- {
- fprintf(stderr, "nrlsmf --cli: failed to send command to instance \"%s\"\n", instance);
- smfPipe.Close();
- listenPipe.Close();
- return 1;
- }
+ if (CliExpectsReply(cmd) && !startedServer)
+ {
+ // Status replies are sent on server_pipe, not back to the requester.
+ // Register this process as the (temporary) controller so we receive them.
+ char startMsg[128];
+ snprintf(startMsg, sizeof(startMsg), "smfServerStart %s", listenName);
+ unsigned int numBytes = (unsigned int)strlen(startMsg) + 1;
+ if (!smfPipe.Send(startMsg, numBytes))
+ {
+ fprintf(stderr, "nrlsmf --cli: failed to send smfServerStart to instance \"%s\"\n",
+ instance);
+ smfPipe.Close();
+ listenPipe.Close();
+ return 1;
+ }
+ startedServer = true;
+ }
- int exitStatus = 0;
- if (wantsReply)
- {
- char reply[8192];
- unsigned int replyLen = sizeof(reply);
- if (!CliRecvWithTimeout(listenPipe, reply, replyLen, 2000))
+ unsigned int numBytes = (unsigned int)strlen(cmd) + 1;
+ if (!smfPipe.Send(cmd, numBytes))
{
- fprintf(stderr, "nrlsmf --cli: timed out waiting for reply to \"%s\"\n", cmd);
- exitStatus = 1;
+ fprintf(stderr, "nrlsmf --cli: failed to send command to instance \"%s\"\n", instance);
+ smfPipe.Close();
+ listenPipe.Close();
+ return 1;
}
- else
+
+ if (CliExpectsReply(cmd))
{
+ char reply[8192];
+ unsigned int replyLen = sizeof(reply);
+ if (!CliRecvWithTimeout(listenPipe, reply, replyLen, 2000))
+ {
+ fprintf(stderr, "nrlsmf --cli: timed out waiting for reply to \"%s\"\n", cmd);
+ smfPipe.Close();
+ listenPipe.Close();
+ return 1;
+ }
fwrite(reply, 1, replyLen, stdout);
if ((0 == replyLen) || ('\n' != reply[replyLen - 1]))
fputc('\n', stdout);
@@ -1405,7 +1665,7 @@ int SmfApp::RunControlClient(int argc, const char*const* argv)
smfPipe.Close();
listenPipe.Close();
- return exitStatus;
+ return 0;
} // end SmfApp::RunControlClient()
bool SmfApp::OnStartup(int argc, const char*const* argv)
@@ -1636,6 +1896,14 @@ void SmfApp::OnShutdown()
if (control_pipe.IsOpen()) control_pipe.Close();
if (server_pipe.IsOpen()) server_pipe.Close();
if (underlay_group_socket.IsOpen()) underlay_group_socket.Close();
+ underlay_join_groups.Destroy();
+ if (NULL != underlay_gre_cap)
+ {
+ underlay_gre_cap->Close();
+ delete underlay_gre_cap;
+ underlay_gre_cap = NULL;
+ }
+ underlay_gre_groups.Destroy();
Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
Smf::Interface* iface;
@@ -1885,6 +2153,13 @@ bool SmfApp::OnCommand(const char* cmd, const char* val)
{
PLOG(PL_DEBUG,"Setup to pull VRF data from FRR\n");
smf.SetWithFRR(true);
+#ifdef ELASTIC_MCAST
+ // Command-line with-frr runs before OnStartup() opens the
+ // controller. A runtime --cli "with-frr" must start FRR polling
+ // on the already-open controller.
+ if (igmp_controller.IsOpen())
+ igmp_controller.EnableFrrPolling();
+#endif // ELASTIC_MCAST
return true;
}
else if (!strncmp("ipv6", cmd, len))
@@ -2225,6 +2500,23 @@ bool SmfApp::OnCommand(const char* cmd, const char* val)
PLOG(PL_ERROR, "SmfApp::OnCommand(elastic) error: unable to retrieve interface name\n");
return false;
}
+ // mGRE PF_PACKET does not deliver 224.0.0.55 unless the
+ // tunnel joins it. EM_ACK/ADV/NACK all use that group.
+ if (iface->IsGRE())
+ {
+ if (!underlay_group_socket.IsOpen() && !underlay_group_socket.Open())
+ {
+ PLOG(PL_ERROR, "SmfApp::OnCommand(elastic) error: unable to open socket to join %s on %s\n",
+ ElasticAck::ELASTIC_ADDR.GetHostString(), ifaceName);
+ return false;
+ }
+ if (!underlay_group_socket.JoinGroup(ElasticAck::ELASTIC_ADDR, ifaceName))
+ {
+ PLOG(PL_ERROR, "SmfApp::OnCommand(elastic) error: join %s on %s failed\n",
+ ElasticAck::ELASTIC_ADDR.GetHostString(), ifaceName);
+ return false;
+ }
+ }
ProtoAddressList groupList;
if (!ProtoNet::GetGroupMemberships(ifaceName, ProtoAddress::IPv4, groupList))
{
@@ -3081,19 +3373,36 @@ bool SmfApp::OnCommand(const char* cmd, const char* val)
return false;
}
addrText = tk.GetNextItem();
+ bool dynamicLearn = false;
ProtoAddress remoteAddr;
- if ((NULL != addrText) && !remoteAddr.ResolveFromString(addrText))
+ if ((NULL != addrText) && (0 == strcmp(addrText, "dynamic")))
+ {
+ dynamicLearn = true;
+ }
+ else if ((NULL != addrText) && !remoteAddr.ResolveFromString(addrText))
{
PLOG(PL_ERROR, "OnCommand(%s) error: invalid remote address \"%s\"\n", cmd, addrText);
return false;
}
TRACE("mapping GRE local:%s", localAddr.GetHostString());
- TRACE(" remote:%s\n", remoteAddr.GetHostString());
+ if (dynamicLearn)
+ TRACE(" remote:dynamic\n");
+ else
+ TRACE(" remote:%s\n", remoteAddr.GetHostString());
if (map)
{
- if (remoteAddr.IsValid())
+ if (dynamicLearn)
+ {
+ iface->SetTunnelLocalAddress(localAddr);
+ iface->SetTunnelLearnDynamic(true);
+ DumpLearnedTunnelNeighbors(*iface);
+ }
+ else if (remoteAddr.IsValid())
{
- // It's a tunnel interface mapping
+ // Multiple map commands with different remotes record multiple
+ // inject destinations for overlay multicast (unicast peers
+ // and/or an underlay multicast group).
+ // 0.0.0.0 is the kernel wildcard remote (not a send dest).
iface->SetTunnelLocalAddress(localAddr);
iface->SetTunnelRemoteAddress(remoteAddr);
if (!smf.AddTunnelInfo(iface->GetIndex(), localAddr, remoteAddr))
@@ -3101,6 +3410,14 @@ bool SmfApp::OnCommand(const char* cmd, const char* val)
PLOG(PL_ERROR, "OnCommand(%s) error: Smf::AddTunnelInfo() failed\n", cmd);
return false;
}
+ if (remoteAddr.IsMulticast() &&
+ GreDeviceIsUnicastMgre(*iface) &&
+ !MaybeEnableUnderlayGreDemux(remoteAddr))
+ {
+ PLOG(PL_ERROR, "OnCommand(%s) error: underlay GRE demux failed for %s\n",
+ cmd, remoteAddr.GetHostString());
+ return false;
+ }
}
else if (!smf.AddOwnAddress(localAddr, iface->GetIndex()))
{
@@ -3108,11 +3425,24 @@ bool SmfApp::OnCommand(const char* cmd, const char* val)
return false;
}
}
+ else if (dynamicLearn)
+ {
+ iface->SetTunnelLearnDynamic(false);
+ ClearLearnedTunnelNeighbors(*iface);
+ }
else if (remoteAddr.IsValid())
{
- iface->SetTunnelLocalAddress(PROTO_ADDR_NONE);
- iface->SetTunnelRemoteAddress(PROTO_ADDR_NONE);
smf.RemoveTunnelInfo(localAddr, remoteAddr);
+ if (remoteAddr.IsMulticast())
+ {
+ underlay_gre_groups.Remove(remoteAddr);
+ if (underlay_gre_groups.IsEmpty() && (NULL != underlay_gre_cap))
+ {
+ underlay_gre_cap->Close();
+ delete underlay_gre_cap;
+ underlay_gre_cap = NULL;
+ }
+ }
}
else
{
@@ -4124,7 +4454,9 @@ bool SmfApp::JoinUnderlayGroup(const ProtoAddress& groupAddr, const char* ifaceN
PLOG(PL_ERROR, "SmfApp::JoinUnderlayGroup() error: group join failed!");
return false;
}
- return true;
+ unsigned int ifaceIndex = ProtoNet::GetInterfaceIndex(ifaceName);
+ underlay_join_groups.Insert(groupAddr, INT2VOIDP(ifaceIndex));
+ return MaybeEnableUnderlayGreDemux(groupAddr);
} // end SmfApp::JoinUnderlayGroup()
bool SmfApp::LeaveUnderlayGroup(const ProtoAddress& groupAddr, const char* ifaceName)
@@ -4135,9 +4467,278 @@ bool SmfApp::LeaveUnderlayGroup(const ProtoAddress& groupAddr, const char* iface
PLOG(PL_ERROR, "SmfApp::LeaveUnderlayGroup() error: group leave failed!");
return false;
}
+ underlay_join_groups.Remove(groupAddr);
+ underlay_gre_groups.Remove(groupAddr);
+ if (underlay_gre_groups.IsEmpty() && (NULL != underlay_gre_cap))
+ {
+ underlay_gre_cap->Close();
+ delete underlay_gre_cap;
+ underlay_gre_cap = NULL;
+ }
return true;
} // end SmfApp::LeaveUnderlayGroup()
+bool SmfApp::GreDeviceIsUnicastMgre(Smf::Interface& iface)
+{
+ InterfaceMechanism* mech = static_cast(iface.GetExtension());
+ if ((NULL == mech) || (NULL == mech->GetPrincipalElement()))
+ return false;
+ const ProtoAddress& kernRemote =
+ mech->GetPrincipalElement()->GetProtoCap().GetTunnelRemoteAddr();
+ if (!kernRemote.IsValid() || kernRemote.IsUnspecified())
+ return true;
+ return kernRemote.IsUnicast();
+} // end SmfApp::GreDeviceIsUnicastMgre()
+
+bool SmfApp::MaybeEnableUnderlayGreDemux(const ProtoAddress& groupAddr)
+{
+ if (!underlay_join_groups.Contains(groupAddr))
+ return true;
+ unsigned int greIndex = smf.FindMappedIndexForRemote(groupAddr);
+ if (0 == greIndex)
+ return true;
+ Smf::Interface* greIface = smf.GetInterface(greIndex);
+ if ((NULL == greIface) || !GreDeviceIsUnicastMgre(*greIface))
+ return true;
+ unsigned int ifIndex =
+ (unsigned int)(uintptr_t)underlay_join_groups.GetUserData(groupAddr);
+ char ifaceName[Smf::IF_NAME_MAX + 1];
+ ifaceName[Smf::IF_NAME_MAX] = '\0';
+ if (0 == ProtoNet::GetInterfaceName(ifIndex, ifaceName, Smf::IF_NAME_MAX))
+ {
+ PLOG(PL_ERROR, "SmfApp::MaybeEnableUnderlayGreDemux() error: no name for ifIndex %u\n",
+ ifIndex);
+ return false;
+ }
+ return EnableUnderlayGreDemux(groupAddr, ifaceName);
+} // end SmfApp::MaybeEnableUnderlayGreDemux()
+
+bool SmfApp::EnableUnderlayGreDemux(const ProtoAddress& groupAddr, const char* ifaceName)
+{
+ underlay_gre_groups.Insert(groupAddr);
+ if (NULL != underlay_gre_cap)
+ return true;
+ underlay_gre_cap = ProtoCap::Create();
+ if (NULL == underlay_gre_cap)
+ {
+ PLOG(PL_ERROR, "SmfApp::EnableUnderlayGreDemux() ProtoCap::Create() error\n");
+ return false;
+ }
+ underlay_gre_cap->SetListener(this, &SmfApp::OnUnderlayGreCapture);
+ underlay_gre_cap->SetNotifier(static_cast(&dispatcher));
+ if (!underlay_gre_cap->Open(ifaceName))
+ {
+ PLOG(PL_ERROR, "SmfApp::EnableUnderlayGreDemux() ProtoCap::Open(%s) error\n", ifaceName);
+ delete underlay_gre_cap;
+ underlay_gre_cap = NULL;
+ return false;
+ }
+ if (!underlay_gre_cap->StartInputNotification())
+ {
+ PLOG(PL_ERROR, "SmfApp::EnableUnderlayGreDemux() StartInputNotification() error\n");
+ underlay_gre_cap->Close();
+ delete underlay_gre_cap;
+ underlay_gre_cap = NULL;
+ return false;
+ }
+ return true;
+} // end SmfApp::EnableUnderlayGreDemux()
+
+void SmfApp::OnUnderlayGreCapture(ProtoChannel& theChannel,
+ ProtoChannel::Notification notifyType)
+{
+ if (ProtoChannel::NOTIFY_INPUT != notifyType)
+ return;
+ ProtoCap& cap = static_cast(theChannel);
+ UINT32 alignedBuffer[BUFFER_MAX/sizeof(UINT32)];
+ UINT16* ethBuffer = ((UINT16*)(alignedBuffer + 256)) + 1;
+ const unsigned int ETHER_BYTES_MAX = (BUFFER_MAX - 256 * sizeof(UINT32) - 2);
+ for (;;)
+ {
+ unsigned int numBytes = ETHER_BYTES_MAX;
+ ProtoCap::Direction direction;
+ if (!cap.Recv((char*)ethBuffer, numBytes, &direction))
+ {
+ PLOG(PL_ERROR, "SmfApp::OnUnderlayGreCapture() ProtoCap::Recv() error\n");
+ break;
+ }
+ if (0 == numBytes)
+ break;
+ if (ProtoCap::INBOUND != direction)
+ continue;
+
+ ProtoPktETH ethPkt((UINT32*)ethBuffer, ETHER_BYTES_MAX);
+ if (!ethPkt.InitFromBuffer(numBytes) || (ProtoPktETH::IP != ethPkt.GetType()))
+ continue;
+ ProtoPktIP ipPkt;
+ if (!ipPkt.InitFromBuffer(ethPkt.GetPayloadLength(),
+ ethPkt.AccessPayload(),
+ ethPkt.GetPayloadLength()) ||
+ (4 != ipPkt.GetVersion()))
+ continue;
+ ProtoPktIPv4 ip4(ipPkt);
+ if (ProtoPktIP::GRE != ip4.GetProtocol())
+ continue;
+ ProtoAddress dst;
+ ip4.GetDstAddr(dst);
+ if (!underlay_gre_groups.Contains(dst))
+ continue;
+
+ unsigned int ipHdrLen = ip4.GetHeaderLength();
+ unsigned int greBytes = ethPkt.GetPayloadLength();
+ if (greBytes <= ipHdrLen)
+ continue;
+ greBytes -= ipHdrLen;
+ const UINT8* grePtr = (const UINT8*)ethPkt.AccessPayload() + ipHdrLen;
+ if (greBytes < 4)
+ continue;
+ unsigned int greHdrLen = 4;
+ if (0 != (grePtr[0] & 0x80)) greHdrLen += 4; // checksum
+ if (0 != (grePtr[0] & 0x20)) greHdrLen += 4; // key
+ if (0 != (grePtr[0] & 0x10)) greHdrLen += 4; // sequence
+ if (greBytes <= greHdrLen)
+ continue;
+ unsigned int innerLen = greBytes - greHdrLen;
+ const UINT8* inner = grePtr + greHdrLen;
+
+ unsigned int greIndex = smf.FindMappedIndexForRemote(dst);
+ if (0 == greIndex)
+ continue;
+ Smf::Interface* greIface = smf.GetInterface(greIndex);
+ if (NULL == greIface)
+ continue;
+ InterfaceMechanism* mech = static_cast(greIface->GetExtension());
+ if ((NULL == mech) || (NULL == mech->GetPrincipalElement()))
+ continue;
+ ProtoCap& greCap = mech->GetPrincipalElement()->GetProtoCap();
+
+ UINT8 innerCopy[FRAME_SIZE_MAX];
+ if (innerLen > FRAME_SIZE_MAX)
+ continue;
+ memcpy(innerCopy, inner, innerLen);
+ memcpy((char*)ethBuffer + 14, innerCopy, innerLen);
+ ProtoPktETH outEth;
+ outEth.InitIntoBuffer(ethBuffer, 14 + innerLen);
+ outEth.SetType(ProtoPktETH::IP);
+ outEth.SetPayloadLength(innerLen);
+ HandleInboundPacket(alignedBuffer, 14 + innerLen, greCap);
+ }
+} // end SmfApp::OnUnderlayGreCapture()
+
+static bool NeighStateUsable(unsigned short state)
+{
+#ifdef NUD_PERMANENT
+ return (0 != (state & (NUD_PERMANENT | NUD_REACHABLE | NUD_STALE |
+ NUD_DELAY | NUD_PROBE | NUD_NOARP)));
+#else
+ return (0 != state);
+#endif
+}
+
+bool SmfApp::OnNeighborDump(unsigned int ifIndex,
+ const ProtoAddress& dst,
+ const ProtoAddress& lladdr,
+ unsigned short ndmState,
+ void* userData)
+{
+ SmfApp* app = static_cast(userData);
+ Smf::Interface* iface = app->smf.GetInterface(ifIndex);
+ if ((NULL == iface) || !iface->GetTunnelLearnDynamic())
+ return true;
+ app->UpdateLearnedTunnelPeer(*iface, dst, lladdr, ndmState, false);
+ return true;
+} // end SmfApp::OnNeighborDump()
+
+void SmfApp::DumpLearnedTunnelNeighbors(Smf::Interface& iface)
+{
+ if (!ProtoNet::GetInterfaceNeighbors(iface.GetIndex(), OnNeighborDump, this))
+ PLOG(PL_WARN, "SmfApp::DumpLearnedTunnelNeighbors() warning: neighbor dump failed for %s\n",
+ iface.GetNameStr());
+} // end SmfApp::DumpLearnedTunnelNeighbors()
+
+void SmfApp::ClearLearnedTunnelNeighbors(Smf::Interface& iface)
+{
+ const ProtoAddress& local = iface.GetTunnelLocalAddress();
+ ProtoAddress overlay;
+ ProtoAddressList::Iterator it(iface.AccessLearnedOverlays());
+ while (it.GetNextAddress(overlay))
+ {
+ const ProtoAddress* underlay =
+ static_cast(iface.AccessLearnedOverlays().GetUserData(overlay));
+ if ((NULL != underlay) && local.IsValid())
+ smf.ClearTunnelSource(local, *underlay, false, true);
+ }
+ iface.ClearLearnedOverlays();
+} // end SmfApp::ClearLearnedTunnelNeighbors()
+
+void SmfApp::UpdateLearnedTunnelPeer(Smf::Interface& iface,
+ const ProtoAddress& overlay,
+ const ProtoAddress& underlay,
+ unsigned short ndmState,
+ bool deleted)
+{
+ if (!overlay.IsValid() || !overlay.IsUnicast())
+ return;
+ const ProtoAddress& local = iface.GetTunnelLocalAddress();
+ if (!local.IsValid())
+ return;
+
+ const bool usable = !deleted &&
+ underlay.IsValid() && underlay.IsUnicast() &&
+ !underlay.HostIsEqual(local) &&
+ NeighStateUsable(ndmState);
+
+ ProtoAddressList& learned = iface.AccessLearnedOverlays();
+ const ProtoAddress* oldUnderlay =
+ static_cast(learned.GetUserData(overlay));
+
+ if (!usable)
+ {
+ if (NULL != oldUnderlay)
+ {
+ smf.ClearTunnelSource(local, *oldUnderlay, false, true);
+ delete const_cast(oldUnderlay);
+ learned.Remove(overlay);
+ }
+ return;
+ }
+
+ if ((NULL != oldUnderlay) && oldUnderlay->HostIsEqual(underlay))
+ return; // already have this mapping
+
+ if (NULL != oldUnderlay)
+ {
+ smf.ClearTunnelSource(local, *oldUnderlay, false, true);
+ delete const_cast(oldUnderlay);
+ learned.Remove(overlay);
+ }
+
+ ProtoAddress* stored = new ProtoAddress(underlay);
+ if (NULL == stored)
+ {
+ PLOG(PL_ERROR, "SmfApp::UpdateLearnedTunnelPeer() new ProtoAddress error: %s\n", GetErrorString());
+ return;
+ }
+ if (!learned.Insert(overlay, stored))
+ {
+ PLOG(PL_ERROR, "SmfApp::UpdateLearnedTunnelPeer() error inserting learned overlay %s\n",
+ overlay.GetHostString());
+ delete stored;
+ return;
+ }
+ if (!smf.AddTunnelInfo(iface.GetIndex(), local, underlay, false, true))
+ {
+ PLOG(PL_ERROR, "SmfApp::UpdateLearnedTunnelPeer() AddTunnelInfo() failed for %s\n",
+ underlay.GetHostString());
+ learned.Remove(overlay);
+ delete stored;
+ return;
+ }
+ PLOG(PL_DEBUG, "SmfApp::UpdateLearnedTunnelPeer() %s overlay %s",
+ iface.GetNameStr(), overlay.GetHostString());
+ PLOG(PL_DEBUG, " -> underlay %s\n", underlay.GetHostString());
+} // end SmfApp::UpdateLearnedTunnelPeer()
+
// This method gets (creates as needed) and configures an interface group
Smf::InterfaceGroup* SmfApp::GetInterfaceGroup(const char* groupName,
Smf::Mode mode,
@@ -4471,7 +5072,7 @@ Smf::Interface* SmfApp::GetInterface(const char* ifName, unsigned int ifIndex)
InterfaceMechanism* mech = static_cast(iface->GetExtension());
if (NULL == mech)
{
- if (NULL == (mech = new InterfaceMechanism(*iface, pkt_pool)))
+ if (NULL == (mech = new InterfaceMechanism(*iface, pkt_pool, smf)))
{
PLOG(PL_ERROR, "SmfApp::GetInterface(): new InterfaceMechanism error: %s\n", GetErrorString());
smf.RemoveInterface(ifIndex);
@@ -4523,7 +5124,7 @@ Smf::Interface* SmfApp::GetInterface(const char* ifName, unsigned int ifIndex)
{
iface->SetTunnelLocalAddress(localAddr);
iface->SetTunnelRemoteAddress(remoteAddr);
- smf.AddTunnelInfo(ifIndex, localAddr, remoteAddr);
+ smf.AddTunnelInfo(ifIndex, localAddr, remoteAddr, false);
}
else
{
@@ -5403,7 +6004,7 @@ Smf::Interface* SmfApp::AddDevice(const char* vifName, const char* ifaceNameAndF
{
iface->SetTunnelLocalAddress(localAddr);
iface->SetTunnelRemoteAddress(remoteAddr);
- smf.AddTunnelInfo(vifIndex, localAddr, remoteAddr);
+ smf.AddTunnelInfo(vifIndex, localAddr, remoteAddr, false);
}
else
{
@@ -5474,7 +6075,7 @@ Smf::Interface* SmfApp::CreateDevice(const char* vifName)
}
// Create InterfaceMechanism to associate vif device
- InterfaceMechanism* mech = new InterfaceMechanism(*iface, pkt_pool);
+ InterfaceMechanism* mech = new InterfaceMechanism(*iface, pkt_pool, smf);
if (NULL == mech)
{
PLOG(PL_ERROR, "SmfApp::CreateDevice() new InterfaceMechanism error: %s\n", GetErrorString());
@@ -5740,42 +6341,745 @@ bool SmfApp::AssignAddresses(const char* ifaceName, unsigned int ifaceIndex, con
return true;
} // end SmfApp::AssignAddresses()
-/* These are the messages that come in through the server socket
- * "-jsonInfo", "Returns string with group names and interfaces, json formatted to unix socket",
- * "-jsonStats", "Return stats for everything in json format to unix socket",
- * "-jsonVersion", "Return version in json format to unix socket", Not in CLI
- * "-ping", "Ping/Heartbeat returns 'pong' if nrlsmf is running", // should this be for a specific group?
- * "-stats", "Return stats for everything to unix socket",
- * "-info", "Returns string with group names and interfaces"
- */
-void SmfApp::OnControlMsg(ProtoSocket& thePipe, ProtoSocket::Event theEvent)
+bool SmfApp::ControlReply(const char* data, unsigned int numBytes)
{
- if (ProtoSocket::RECV == theEvent)
+ if (!server_pipe.IsOpen())
{
- char buffer[8192];
- unsigned int len = 8191;
- if (thePipe.Recv(buffer, len))
- {
- // trim trailing white space if present
- char *end = buffer + len - 1;
- while(end > buffer && isspace((unsigned char)*end)) end--;
- end[1] = '\0';
- len = strlen(buffer);
+ fprintf(stderr, "Server pipe is NOT open\n");
+ PLOG(PL_WARN, "SmfApp::ControlReply() server pipe is not open\n");
+ return false;
+ }
+ if (!server_pipe.Send(data, numBytes))
+ {
+ PLOG(PL_ERROR, "SmfApp::ControlReply() error sending %u byte reply\n", numBytes);
+ return false;
+ }
+ return true;
+}
- char* arg = NULL;
+bool SmfApp::ControlReply(const std::string& s)
+{
+ unsigned int n = (unsigned int)s.size();
+ return ControlReply(s.c_str(), n);
+}
- for (unsigned int i = 0; i < len; i++)
- {
- if ('\0' == buffer[i])
- {
- break;
- }
- else if (isspace(buffer[i]))
- {
- buffer[i] = '\0';
- arg = buffer+i+1; // ending arg may be empty if buffer ends in " \n" so trim when we first get it
- break;
- }
+void SmfApp::ReplyVersion(bool json)
+{
+ if (json)
+ {
+ ServerSend("jsonVersion", _SMF_VERSION);
+ return;
+ }
+ char buf[128];
+ snprintf(buf, sizeof(buf), "smf version: %s\n", _SMF_VERSION);
+ ControlReply(std::string(buf));
+}
+
+void SmfApp::ReplyStats(bool json)
+{
+ std::ostringstream ss;
+ Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
+ Smf::Interface* nextIface;
+ if (json)
+ {
+ ss << "[";
+ bool comma = false;
+ while (NULL != (nextIface = iterator.GetNextItem()))
+ {
+ ss << (comma ? "," : "") << "{";
+ ss << "\"interface\":\"" << nextIface->GetNameStr() << "\",";
+ ss << "\"flows\":\"" << nextIface->GetFlowCount() << "\",";
+ ss << "\"recv\":\"" << nextIface->GetRecvCount() << "\",";
+ ss << "\"mrcv\":\"" << nextIface->GetMcastCount() << "\",";
+ ss << "\"sent\":\"" << nextIface->GetSentCount() << "\",";
+ ss << "\"retr\":\"" << nextIface->GetRetransmissionCount() << "\",";
+ ss << "\"fwd\":\"" << nextIface->GetForwardCount() << "\",";
+ ss << "\"dups\":\"" << nextIface->GetDuplicateCount() << "\",";
+ ss << "\"asym\":\"" << nextIface->GetAsymCount() << "\",";
+ ss << "\"queue\":\"" << nextIface->GetQueueLength() << "\"";
+ ss << "}";
+ comma = true;
+ }
+ ss << "]\n";
+ }
+ else
+ {
+ ss << "Interface Flows Receives MReceives Sends ReXmits Forwards Duplicates Asyms QueueLen\n";
+ ss << "---------------- ---------- ---------- ---------- ---------- ---------- ---------- ---------- ---------- ----------\n";
+ while (NULL != (nextIface = iterator.GetNextItem()))
+ {
+ ss << std::left << std::setw(16) << nextIface->GetNameStr() << " ";
+ ss << std::right << std::setw(10) << nextIface->GetFlowCount() << " ";
+ ss << std::setw(10) << nextIface->GetRecvCount() << " ";
+ ss << std::setw(10) << nextIface->GetMcastCount() << " ";
+ ss << std::setw(10) << nextIface->GetSentCount() << " ";
+ ss << std::setw(10) << nextIface->GetRetransmissionCount() << " ";
+ ss << std::setw(10) << nextIface->GetForwardCount() << " ";
+ ss << std::setw(10) << nextIface->GetDuplicateCount() << " ";
+ ss << std::setw(10) << nextIface->GetAsymCount() << " ";
+ ss << std::setw(10) << nextIface->GetQueueLength() << "\n";
+ }
+ }
+ ControlReply(ss.str());
+}
+
+void SmfApp::ReplyInfo(bool json)
+{
+ Smf::InterfaceGroupList::Iterator grouperator(smf.AccessInterfaceGroupList());
+ Smf::InterfaceGroup* group;
+ std::ostringstream ss;
+ if (json)
+ {
+ bool first = true;
+ std::string spot;
+ ss << "[";
+ while (NULL != (group = grouperator.GetNextItem()))
+ {
+ char ifaceName[Smf::IF_NAME_MAX + 1];
+ ifaceName[Smf::IF_NAME_MAX] = '\0';
+ Smf::InterfaceGroup::Iterator ifacerator(*group);
+ Smf::Interface* iface;
+
+ spot = first ? "" : ",";
+ first = false;
+ ss << spot << "{\"GroupName\": \"" << group->GetName() << "\",";
+ ss << "\"GroupType\": \"" << (group->IsTemplateGroup() ? "Template" : "Regular") << "\",";
+ std::string relayType;
+ switch (group->GetRelayType())
+ {
+ case Smf::INVALID: relayType="Invalid"; break;
+ case Smf::CF: relayType="cf"; break;
+ case Smf::S_MPR: relayType="s_mpr"; break;
+ case Smf::E_CDS: relayType="e_cds"; break;
+ case Smf::MPR_CDS: relayType="mpr_cds"; break;
+ case Smf::NS_MPR: relayType="ns_mpr"; break;
+ }
+ ss << "\"RelayType\": \"" << relayType << "\",";
+ switch (group->GetForwardingMode())
+ {
+ case Smf::PUSH: relayType="Push"; break;
+ case Smf::MERGE: relayType="Merge"; break;
+ case Smf::RELAY: relayType="Relay"; break;
+ }
+ ss << "\"ForwardingMode\": \"" << relayType << "\",";
+ ss << "\"Interfaces\": [";
+ bool firstInterface = true;
+ while (NULL != (iface = ifacerator.GetNextInterface()))
+ {
+ ProtoNet::GetInterfaceName(iface->GetIndex(), ifaceName, Smf::IF_NAME_MAX);
+ spot = firstInterface ? "" : ",";
+ ss << spot << "\""<< ifaceName << "\"";
+ firstInterface = false;
+ }
+ ss << "]";
+ if (group->GetElasticMulticast())
+ ss << ", \"Elastic\" : true";
+ if (group->GetAdaptiveRouting())
+ ss << ", \"Adaptive\" : true";
+ ss << "}";
+ }
+ ss << "]\n";
+ }
+ else
+ {
+ ss << "GroupName GroupType RelayType ForwardingMode Interfaces\n";
+ ss << "-------------------- --------- --------- -------------- ----------\n";
+ while (NULL != (group = grouperator.GetNextItem()))
+ {
+ char ifaceName[Smf::IF_NAME_MAX + 1];
+ ifaceName[Smf::IF_NAME_MAX] = '\0';
+ Smf::InterfaceGroup::Iterator ifacerator(*group);
+ Smf::Interface* iface;
+
+ ss << std::left << std::setw(21) << group->GetName();
+ ss << std::setw(9) << (group->IsTemplateGroup() ? "Template" : "Regular") << " ";
+ std::string relayType;
+ switch (group->GetRelayType())
+ {
+ case Smf::INVALID: relayType="Invalid"; break;
+ case Smf::CF: relayType="cf"; break;
+ case Smf::S_MPR: relayType="s_mpr"; break;
+ case Smf::E_CDS: relayType="e_cds"; break;
+ case Smf::MPR_CDS: relayType="mpr_cds"; break;
+ case Smf::NS_MPR: relayType="ns_mpr"; break;
+ }
+ ss << std::setw(9) << relayType << " ";
+ switch (group->GetForwardingMode())
+ {
+ case Smf::PUSH: relayType="Push"; break;
+ case Smf::MERGE: relayType="Merge"; break;
+ case Smf::RELAY: relayType="Relay"; break;
+ }
+ ss << std::setw(14) << relayType << " ";
+ bool firstInterface = true;
+ while (NULL != (iface = ifacerator.GetNextInterface()))
+ {
+ ProtoNet::GetInterfaceName(iface->GetIndex(), ifaceName, Smf::IF_NAME_MAX);
+ ss << ( firstInterface ? "" : ",") << ifaceName;
+ firstInterface = false;
+ }
+ if (group->GetElasticMulticast())
+ ss << ", Elastic";
+ if (group->GetAdaptiveRouting())
+ ss << ", Adaptive";
+ ss << "\n";
+ }
+ ss << "\n";
+ }
+ ControlReply(ss.str());
+}
+
+void SmfApp::ReplyInterfaces(bool json)
+{
+ std::ostringstream ss;
+ Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
+ Smf::Interface* nextIface;
+ if (json)
+ {
+ bool comma = false;
+ ss << "[";
+ while (NULL != (nextIface = iterator.GetNextItem()))
+ {
+ ss << (comma ? "," : "") << "{";
+ ss << "\"Interface\" : \"" << nextIface->GetNameStr() << "\",";
+ ss << "\"FwdMethod\" : \"";
+#ifdef ELASTIC_MCAST
+ if (nextIface->GetElasticMulticast()) {
+ if (mcast_controller.GetDefaultForwardingStatus() == MulticastFIB::HYBRID)
+ ss << "Advertise";
+ else
+ ss << "Elastic";
+ } else ss << "Flood";
+#else
+ ss << "Flood";
+#endif // ELASTIC_MCAST
+ ss << "\",";
+ ss << "\"Flags\" : \"";
+ if (nextIface->IsLayered()) ss << "L";
+ if (nextIface->IsTunnel()) ss << "T";
+ if (nextIface->IsIgmpProxy()) ss << "I";
+ InterfaceMechanism* mech = static_cast(nextIface->GetExtension());
+ if ((NULL != mech) && mech->IsShadowing()) ss << "S";
+#ifdef ELASTIC_MCAST
+ if (nextIface->IsManaged()) ss << "M";
+#endif // ELASTIC_MCAST
+ ss << "\"";
+#ifdef ELASTIC_MCAST
+ ss << ", \"Managed\" : " << (nextIface->IsManaged() ? "true" : "false");
+#endif // ELASTIC_MCAST
+ ss << "}";
+ comma = true;
+ }
+ ss << "]\n";
+ }
+ else
+ {
+ ss << "Flags: L = Layered, T = Tunnel, I = IGMP Proxy, S = Shadowing";
+#ifdef ELASTIC_MCAST
+ ss << ", M = Managed";
+#endif // ELASTIC_MCAST
+ ss << "\n\n";
+ ss << "Interface Fwd Method Flags\n";
+ ss << "---------------- ---------- -----\n";
+ while (NULL != (nextIface = iterator.GetNextItem()))
+ {
+ ss << std::left << std::setw(16) << nextIface->GetNameStr() << " ";
+ ss << std::setw(12);
+#ifdef ELASTIC_MCAST
+ if (nextIface->GetElasticMulticast()) {
+ if (mcast_controller.GetDefaultForwardingStatus() == MulticastFIB::HYBRID)
+ ss << "Advertise";
+ else
+ ss << "Elastic";
+ } else ss << "Flood";
+#else
+ ss << "Flood";
+#endif // ELASTIC_MCAST
+ if (nextIface->IsLayered()) ss << "L";
+ if (nextIface->IsTunnel()) ss << "T";
+ if (nextIface->IsIgmpProxy()) ss << "I";
+ InterfaceMechanism* mech = static_cast(nextIface->GetExtension());
+ if ((NULL != mech) && mech->IsShadowing()) ss << "S";
+#ifdef ELASTIC_MCAST
+ if (nextIface->IsManaged()) ss << "M";
+#endif // ELASTIC_MCAST
+ ss << "\n";
+ }
+ }
+ ControlReply(ss.str());
+}
+
+// C = Config. Neighbor NUD letters (linux/neighbour.h) imply kernel:
+// M Permanent, N NoARP, R Reachable, S Stale, D Delay, P Probe,
+// I Incomplete, F Failed.
+static void FormatTunnelFlags(char* buf, size_t bufLen, bool fromConfig,
+ unsigned short nudState = 0)
+{
+ size_t n = 0;
+ if (fromConfig && (n + 1 < bufLen)) buf[n++] = 'C';
+ char nud = '\0';
+ if (nudState & 0x80) nud = 'M';
+ else if (nudState & 0x40) nud = 'N';
+ else if (nudState & 0x02) nud = 'R';
+ else if (nudState & 0x04) nud = 'S';
+ else if (nudState & 0x08) nud = 'D';
+ else if (nudState & 0x10) nud = 'P';
+ else if (nudState & 0x01) nud = 'I';
+ else if (nudState & 0x20) nud = 'F';
+ if (nud && (n + 1 < bufLen)) buf[n++] = nud;
+ if (bufLen > 0) buf[n] = '\0';
+}
+
+static bool TunnelAddrUnspecified(const ProtoAddress& addr)
+{
+ if (!addr.IsValid())
+ return true;
+ return addr.HostIsEqual(PROTO_ADDR_ANY) || addr.HostIsEqual(PROTO_ADDR_ANY6);
+}
+
+static bool SameHostAddr(const ProtoAddress& a, const ProtoAddress& b)
+{
+ return a.IsValid() && b.IsValid() && a.HostIsEqual(b);
+}
+
+static void FormatAddrOrDash(const ProtoAddress& addr, char* buf, size_t bufLen)
+{
+ if (TunnelAddrUnspecified(addr))
+ strncpy(buf, "-", bufLen);
+ else
+ addr.GetHostString(buf, bufLen);
+ buf[bufLen - 1] = '\0';
+}
+
+static void TunnelOverlayLocal(Smf::Interface* iface, ProtoAddress& overlay)
+{
+ if (NULL == iface)
+ return;
+ const ProtoAddress& ip = iface->GetIpAddress();
+ if (ip.IsValid() && (ProtoAddress::ETH != ip.GetType()))
+ {
+ overlay = ip;
+ return;
+ }
+ ProtoAddressList::Iterator it(iface->AccessAddressList());
+ ProtoAddress addr;
+ while (it.GetNextAddress(addr))
+ {
+ if ((ProtoAddress::IPv4 == addr.GetType()) || (ProtoAddress::IPv6 == addr.GetType()))
+ {
+ overlay = addr;
+ return;
+ }
+ }
+}
+
+void SmfApp::ReplyTunnel(bool json)
+{
+ std::ostringstream ss;
+ Smf::InterfaceInfoTable::Iterator it(smf.AccessInterfaceInfoTable());
+ Smf::InterfaceInfo* info;
+ if (json)
+ {
+ ss << "[";
+ bool comma = false;
+ while (NULL != (info = it.GetNextItem()))
+ {
+ char ifaceName[Smf::IF_NAME_MAX + 1];
+ ifaceName[0] = '\0';
+ ProtoNet::GetInterfaceName(info->GetIndex(), ifaceName, Smf::IF_NAME_MAX);
+ char localStr[64];
+ char remoteStr[64];
+ char overlayStr[64];
+ info->GetLocalAddress().GetHostString(localStr, sizeof(localStr));
+ FormatAddrOrDash(info->GetRemoteAddress(), remoteStr, sizeof(remoteStr));
+ ProtoAddress overlay;
+ if (info->GetRemoteAddress().IsValid())
+ TunnelOverlayLocal(smf.GetInterface(info->GetIndex()), overlay);
+ FormatAddrOrDash(overlay, overlayStr, sizeof(overlayStr));
+ char flags[8];
+ FormatTunnelFlags(flags, sizeof(flags), info->FromConfig());
+ ss << (comma ? "," : "") << "{";
+ ss << "\"Interface\":\"" << ifaceName << "\",";
+ ss << "\"Local\":\"" << localStr << "\",";
+ ss << "\"Remote\":\"" << remoteStr << "\",";
+ ss << "\"IP\":\"" << overlayStr << "\",";
+ ss << "\"Flags\":\"" << flags << "\"";
+ ss << "}";
+ comma = true;
+ }
+ ss << "]\n";
+ }
+ else
+ {
+ ss << "Flags: C = Config\n";
+ ss << "Local/Remote are underlay tunnel endpoints; IP is the local overlay address.\n\n";
+ ss << "Interface Local Remote IP Flags\n";
+ ss << "---------------- ---------------- ---------------- ---------------- -----\n";
+ while (NULL != (info = it.GetNextItem()))
+ {
+ char ifaceName[Smf::IF_NAME_MAX + 1];
+ ifaceName[0] = '\0';
+ ProtoNet::GetInterfaceName(info->GetIndex(), ifaceName, Smf::IF_NAME_MAX);
+ char localStr[64];
+ char remoteStr[64];
+ char overlayStr[64];
+ info->GetLocalAddress().GetHostString(localStr, sizeof(localStr));
+ FormatAddrOrDash(info->GetRemoteAddress(), remoteStr, sizeof(remoteStr));
+ ProtoAddress overlay;
+ if (info->GetRemoteAddress().IsValid())
+ TunnelOverlayLocal(smf.GetInterface(info->GetIndex()), overlay);
+ FormatAddrOrDash(overlay, overlayStr, sizeof(overlayStr));
+ char flags[8];
+ FormatTunnelFlags(flags, sizeof(flags), info->FromConfig());
+ ss << std::left << std::setw(16) << ifaceName << " ";
+ ss << std::setw(16) << localStr << " ";
+ ss << std::setw(16) << remoteStr << " ";
+ ss << std::setw(16) << overlayStr << " ";
+ ss << flags << "\n";
+ }
+ }
+ ControlReply(ss.str());
+}
+
+struct TunnelNeighRow
+{
+ unsigned int ifIndex;
+ ProtoAddress overlay_remote; // kernel neigh dst
+ ProtoAddress underlay_remote; // kernel neigh lladdr / mapped GRE remote
+ unsigned short state;
+ bool from_config;
+ bool from_kernel;
+};
+
+struct ShowNeighDump
+{
+ std::vector* rows;
+};
+
+static bool ShowNeighHandler(unsigned int ifIndex,
+ const ProtoAddress& dst,
+ const ProtoAddress& lladdr,
+ unsigned short ndmState,
+ void* userData)
+{
+ ShowNeighDump* ctx = static_cast(userData);
+ TunnelNeighRow row;
+ row.ifIndex = ifIndex;
+ row.overlay_remote = dst;
+ row.underlay_remote = lladdr;
+ row.state = ndmState;
+ row.from_config = false;
+ row.from_kernel = true;
+ ctx->rows->push_back(row);
+ return true;
+}
+
+static void MarkTunnelNeighbor(std::vector& rows,
+ unsigned int ifIndex,
+ const ProtoAddress& peer,
+ bool fromConfig,
+ bool fromKernel)
+{
+ bool found = false;
+ for (size_t i = 0; i < rows.size(); i++)
+ {
+ if (rows[i].ifIndex != ifIndex)
+ continue;
+ if (SameHostAddr(rows[i].underlay_remote, peer) || SameHostAddr(rows[i].overlay_remote, peer))
+ {
+ if (fromConfig)
+ rows[i].from_config = true;
+ if (fromKernel)
+ rows[i].from_kernel = true;
+ found = true;
+ }
+ }
+ if (found)
+ return;
+ TunnelNeighRow row;
+ row.ifIndex = ifIndex;
+ row.underlay_remote = peer;
+ row.state = 0;
+ row.from_config = fromConfig;
+ row.from_kernel = fromKernel;
+ rows.push_back(row);
+}
+
+void SmfApp::ReplyTunnelNeighbors(bool json)
+{
+ std::vector rows;
+ ShowNeighDump ctx;
+ ctx.rows = &rows;
+ Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
+ Smf::Interface* iface;
+ while (NULL != (iface = iterator.GetNextItem()))
+ {
+ if (ProtoNet::IFACE_GRE != ProtoNet::GetInterfaceType(iface->GetIndex()))
+ continue;
+ ProtoNet::GetInterfaceNeighbors(iface->GetIndex(), ShowNeighHandler, &ctx);
+ }
+ Smf::InterfaceInfoTable::Iterator it(smf.AccessInterfaceInfoTable());
+ Smf::InterfaceInfo* info;
+ while (NULL != (info = it.GetNextItem()))
+ {
+ const ProtoAddress& remote = info->GetRemoteAddress();
+ if (TunnelAddrUnspecified(remote))
+ continue;
+ // Configured maps and kernel-learned device remotes (P2P, multicast-
+ // underlay). 0.0.0.0 is skipped above; it is not a neighbor.
+ MarkTunnelNeighbor(rows, info->GetIndex(), remote, info->FromConfig(),
+ info->FromKernel());
+ }
+
+ std::ostringstream ss;
+ if (json)
+ {
+ ss << "[";
+ }
+ else
+ {
+ ss << "Flags: C = Config, R = Reachable, S = Stale, D = Delay, P = Probe,\n"
+ << " I = Incomplete, F = Failed, N = NoARP, M = Permanent\n"
+ << "Neighbor IP is the peer overlay address; Remote is the peer underlay.\n\n";
+ ss << "Interface Neighbor IP Remote Flags\n"
+ << "---------------- ---------------- ---------------- -----\n";
+ }
+ bool comma = false;
+ for (size_t i = 0; i < rows.size(); i++)
+ {
+ const TunnelNeighRow& row = rows[i];
+ char ifaceName[Smf::IF_NAME_MAX + 1];
+ ifaceName[0] = '\0';
+ ProtoNet::GetInterfaceName(row.ifIndex, ifaceName, Smf::IF_NAME_MAX);
+ char overlayStr[64];
+ char remoteStr[64];
+ FormatAddrOrDash(row.overlay_remote, overlayStr, sizeof(overlayStr));
+ FormatAddrOrDash(row.underlay_remote, remoteStr, sizeof(remoteStr));
+ char flags[8];
+ FormatTunnelFlags(flags, sizeof(flags), row.from_config, row.state);
+ if (json)
+ {
+ ss << (comma ? "," : "") << "{";
+ ss << "\"Interface\":\"" << ifaceName << "\",";
+ ss << "\"NeighborIP\":\"" << overlayStr << "\",";
+ ss << "\"Remote\":\"" << remoteStr << "\",";
+ ss << "\"Flags\":\"" << flags << "\"";
+ ss << "}";
+ comma = true;
+ }
+ else
+ {
+ ss << std::left << std::setw(16) << ifaceName << " ";
+ ss << std::setw(16) << overlayStr << " ";
+ ss << std::setw(16) << remoteStr << " ";
+ ss << flags << "\n";
+ }
+ }
+ if (json)
+ ss << "]\n";
+ ControlReply(ss.str());
+}
+
+#ifdef ELASTIC_MCAST
+void SmfApp::ReplyGroups(bool json, bool brief, bool details)
+{
+ std::ostringstream ss;
+ mcast_controller.DumpGroups(brief, json, ss, details);
+ ControlReply(ss.str());
+}
+
+void SmfApp::ReplyGroupMemberships(bool json)
+{
+ std::ostringstream ss;
+ mcast_controller.DumpMemberships(json, ss);
+ ControlReply(ss.str());
+}
+
+void SmfApp::ReplyIgmpGroups(bool json)
+{
+ std::ostringstream ss;
+ mcast_controller.DumpManagedGroups(json, ss);
+ ControlReply(ss.str());
+}
+#endif // ELASTIC_MCAST
+
+void SmfApp::OnShowCommand(const char* arg)
+{
+ while ((NULL != arg) && isspace((unsigned char)*arg))
+ arg++;
+ if ((NULL == arg) || ('\0' == *arg) || (0 == strcmp(arg, "?")))
+ {
+ std::ostringstream ss;
+ FormatShowHelp(ss);
+ ControlReply(ss.str());
+ return;
+ }
+
+ char buf[256];
+ strncpy(buf, arg, sizeof(buf) - 1);
+ buf[sizeof(buf) - 1] = '\0';
+
+ const char* topic = NULL;
+ const char* sub = NULL;
+ bool json = false;
+ bool brief = false;
+ bool details = false;
+ char* p = buf;
+ while (*p)
+ {
+ while (isspace((unsigned char)*p))
+ p++;
+ if ('\0' == *p)
+ break;
+ char* start = p;
+ while ((*p != '\0') && !isspace((unsigned char)*p))
+ p++;
+ if (*p)
+ *p++ = '\0';
+ if (NULL == topic)
+ {
+ topic = start;
+ }
+ else if ((NULL == sub) && !json && !brief && !details && IsShowSubcommand(topic, start))
+ {
+ sub = start;
+ }
+ else if (0 == strcmp(start, "json"))
+ {
+ json = true;
+ }
+ else if (json)
+ {
+ ControlReply(std::string("show: 'json' must be the last modifier\n"));
+ return;
+ }
+ else if (0 == strcmp(start, "brief"))
+ {
+ brief = true;
+ }
+ else if ((0 == strcmp(start, "details")) || (0 == strcmp(start, "detail")))
+ {
+ details = true;
+ }
+ else
+ {
+ ControlReply(std::string("show: unknown modifier '") + start + "'\n");
+ return;
+ }
+ }
+
+ if (NULL == topic)
+ {
+ std::ostringstream ss;
+ FormatShowHelp(ss);
+ ControlReply(ss.str());
+ return;
+ }
+ if (brief && details)
+ {
+ ControlReply(std::string("show: 'brief' and 'details' cannot be used together\n"));
+ return;
+ }
+
+ const ShowTopicSpec* spec = FindShowTopic(topic, sub);
+ std::string cmdName = std::string("show ") + topic;
+ if (NULL != sub)
+ cmdName += std::string(" ") + sub;
+ if (NULL == spec)
+ {
+ ControlReply(cmdName + ": unknown command\n");
+ return;
+ }
+ if (json && !spec->json)
+ {
+ ControlReply(cmdName + ": 'json' is not supported\n");
+ return;
+ }
+ if (brief && !spec->brief)
+ {
+ ControlReply(cmdName + ": 'brief' is not supported\n");
+ return;
+ }
+ if (details && !spec->details)
+ {
+ ControlReply(cmdName + ": 'details' is not supported\n");
+ return;
+ }
+
+ if (0 == strcmp(topic, "version"))
+ ReplyVersion(json);
+ else if (0 == strcmp(topic, "statistics"))
+ ReplyStats(json);
+ else if ((0 == strcmp(topic, "interface")) && (NULL != sub) && (0 == strcmp(sub, "grouping")))
+ ReplyInfo(json);
+ else if (0 == strcmp(topic, "interface"))
+ ReplyInterfaces(json);
+ else if ((0 == strcmp(topic, "tunnel")) && (NULL != sub) && (0 == strcmp(sub, "neighbors")))
+ ReplyTunnelNeighbors(json);
+ else if (0 == strcmp(topic, "tunnel"))
+ ReplyTunnel(json);
+ else if ((0 == strcmp(topic, "groups")) && (NULL != sub) && (0 == strcmp(sub, "memberships")))
+ {
+#ifdef ELASTIC_MCAST
+ ReplyGroupMemberships(json);
+#else
+ ControlReply(std::string("show groups memberships is only available in elastic builds\n"));
+#endif // ELASTIC_MCAST
+ }
+ else if ((0 == strcmp(topic, "igmp")) && (NULL != sub) && (0 == strcmp(sub, "groups")))
+ {
+#ifdef ELASTIC_MCAST
+ ReplyIgmpGroups(json);
+#else
+ ControlReply(std::string("show igmp groups is only available in elastic builds\n"));
+#endif // ELASTIC_MCAST
+ }
+ else if (0 == strcmp(topic, "groups"))
+ {
+#ifdef ELASTIC_MCAST
+ // Default listing is DumpGroups(false); brief selects the shorter dump.
+ ReplyGroups(json, brief, details);
+#else
+ ControlReply(std::string("show groups is only available in elastic builds\n"));
+#endif // ELASTIC_MCAST
+ }
+}
+
+/* These are the messages that come in through the server socket
+ * "show [brief|details] [json]" modern CLI status query
+ * e.g. show statistics, show interface, show interface grouping,
+ * show tunnel, show tunnel neighbors
+ * legacy: jsonInfo, jsonStats, jsonVersion, ping, stats, info,
+ * interfaces, interfacesj, groups, groupsj, brfgroups, brfgroupsj
+ */
+void SmfApp::OnControlMsg(ProtoSocket& thePipe, ProtoSocket::Event theEvent)
+{
+ if (ProtoSocket::RECV == theEvent)
+ {
+ char buffer[8192];
+ unsigned int len = 8191;
+ if (thePipe.Recv(buffer, len))
+ {
+ // trim trailing white space if present
+ char *end = buffer + len - 1;
+ while(end > buffer && isspace((unsigned char)*end)) end--;
+ end[1] = '\0';
+ len = strlen(buffer);
+
+ char* arg = NULL;
+
+ for (unsigned int i = 0; i < len; i++)
+ {
+ if ('\0' == buffer[i])
+ {
+ break;
+ }
+ else if (isspace(buffer[i]))
+ {
+ buffer[i] = '\0';
+ arg = buffer+i+1; // ending arg may be empty if buffer ends in " \n" so trim when we first get it
+ break;
+ }
}
// Parse received message from controller and populate our forwarding table.
// If the length of the message is just 1, it is an empty message.
@@ -5933,10 +7237,13 @@ void SmfApp::OnControlMsg(ProtoSocket& thePipe, ProtoSocket::Event theEvent)
}
smf.SetNeighborList(arg, argLen);
}
+ else if (0 == strcmp(cmd, "show"))
+ {
+ OnShowCommand(arg);
+ }
else if (!strncmp("jsonVersion", cmd, len))
{
- ServerSend("jsonVersion", _SMF_VERSION);
- return;
+ ReplyVersion(true);
}
else if (!strncmp("ping", cmd, len)) // just checking that nrlsmf is running, don't care about anything else ...
{
@@ -5951,9 +7258,6 @@ void SmfApp::OnControlMsg(ProtoSocket& thePipe, ProtoSocket::Event theEvent)
}
else
PLOG(PL_DEBUG, "SmfApp::OnCommand(instance) sent heartbeat to smf server\n");
-
- // following line sends json format back, probably not needed
- // ServerSend("ping", "pong");
}
else
{
@@ -5961,332 +7265,38 @@ void SmfApp::OnControlMsg(ProtoSocket& thePipe, ProtoSocket::Event theEvent)
PLOG(PL_WARN, "SmfApp::OnCommand(ping) warning: unable to connect to smfServer\n");
}
}
- else if (!strncmp(cmd, "stats", cmdLen))
- {
- std::ostringstream ss;
- if (server_pipe.IsOpen())
- {
- Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
- Smf::Interface* nextIface;
- ss << "Interface Flows Receives MReceives Sends ReXmits Forwards Duplicates Asyms QueueLen\n";
- ss << "---------------- ---------- ---------- ---------- ---------- ---------- ---------- ---------- ---------- ----------\n";
- while (NULL != (nextIface = iterator.GetNextItem()))
- {
- ss << std::left << std::setw(16) << nextIface->GetNameStr() << " ";
- ss << std::right << std::setw(10) << nextIface->GetFlowCount() << " ";
- ss << std::setw(10) << nextIface->GetRecvCount() << " ";
- ss << std::setw(10) << nextIface->GetMcastCount() << " ";
- ss << std::setw(10) << nextIface->GetSentCount() << " ";
- ss << std::setw(10) << nextIface->GetRetransmissionCount() << " ";
- ss << std::setw(10) << nextIface->GetForwardCount() << " ";
- ss << std::setw(10) << nextIface->GetDuplicateCount() << " ";
- ss << std::setw(10) << nextIface->GetAsymCount() << " ";
- ss << std::setw(10) << nextIface->GetQueueLength() << "\n";
- }
- unsigned int numBytes = ss.str().size();
- if (!server_pipe.Send(ss.str().c_str(), numBytes))
- {
- PLOG(PL_ERROR, "SmfApp::OnCommand(stats) error sending stats to smf server\n");
- return;
- }
- }
- else
- {
- fprintf(stderr, "Server pipe is not open for stats\n");
- return;
- }
+ else if (!strncmp(cmd, "stats", cmdLen))
+ {
+ ReplyStats(false);
}
- else if (!strncmp("jsonInfo", cmd, len)) // just checking groupInfo ...
+ else if (!strncmp("jsonInfo", cmd, len))
{
- Smf::InterfaceGroupList::Iterator grouperator(smf.AccessInterfaceGroupList());
- Smf::InterfaceGroup* group;
- std::ostringstream ss;
- bool first = true;
- std::string spot;
- ss << "[";
- while (NULL != (group = grouperator.GetNextItem()))
- {
- char ifaceName[Smf::IF_NAME_MAX + 1];
- ifaceName[Smf::IF_NAME_MAX] = '\0';
- Smf::InterfaceGroup::Iterator ifacerator(*group);
- Smf::Interface* iface;
-
- // If we don't want to see PUSH groups, uncomment following two lines ...
- // if (Smf::PUSH == group->GetForwardingMode()) // I think we want to skip these ...
- // continue;
- spot = first ? "" : ",";
- first = false;
- ss << spot << "{\"GroupName\": \"" << group->GetName() << "\",";
- ss << "\"GroupType\": \"" << (group->IsTemplateGroup() ? "Template" : "Regular") << "\",";
- std::string relayType;
- switch (group->GetRelayType())
- {
- case Smf::INVALID: relayType="Invalid"; break;
- case Smf::CF: relayType="cf"; break;
- case Smf::S_MPR: relayType="s_mpr"; break;
- case Smf::E_CDS: relayType="e_cds"; break;
- case Smf::MPR_CDS: relayType="mpr_cds"; break;
- case Smf::NS_MPR: relayType="ns_mpr"; break;
- }
- ss << "\"RelayType\": \"" << relayType << "\",";
- switch (group->GetForwardingMode())
- {
- case Smf::PUSH: relayType="Push"; break;
- case Smf::MERGE: relayType="Merge"; break;
- case Smf::RELAY: relayType="Relay"; break;
- }
- ss << "\"ForwardingMode\": \"" << relayType << "\",";
- ss << "\"Interfaces\": [";
- bool firstInterface = true;
- while (NULL != (iface = ifacerator.GetNextInterface()))
- {
- ProtoNet::GetInterfaceName(iface->GetIndex(), ifaceName, Smf::IF_NAME_MAX);
- spot = firstInterface ? "" : ",";
- ss << spot << "\""<< ifaceName << "\"";
- firstInterface = false;
- }
- ss << "]";
- if (group->GetElasticMulticast())
- ss << ", \"Elastic\" : true";
- if (group->GetAdaptiveRouting())
- ss << ", \"Adaptive\" : true";
- ss << "}";
- }
- ss << "]\n";
- if (server_pipe.IsOpen())
- {
- unsigned int numBytes = ss.str().size();
- if (!server_pipe.Send(ss.str().c_str(), numBytes))
- {
- PLOG(PL_ERROR, "SmfApp::OnCommand(jsonInfo) error sending jsonInfo to smf server\n");
- return;
- }
- }
- else
- {
- fprintf(stderr, "Server pipe is NOT open\n");
- PLOG(PL_WARN, "SmfApp::OnCommand(jsonInfo) warning: unable to connect to smfServer\n");
- }
+ ReplyInfo(true);
}
- else if (!strncmp("info", cmd, len)) // just checking groupInfo ...
+ else if (!strncmp("info", cmd, len))
{
- Smf::InterfaceGroupList::Iterator grouperator(smf.AccessInterfaceGroupList());
- Smf::InterfaceGroup* group;
- std::ostringstream ss;
- ss << "";
- ss << "GroupName GroupType RelayType ForwardingMode Interfaces\n";
- ss << "-------------------- --------- --------- -------------- ----------\n";
- while (NULL != (group = grouperator.GetNextItem()))
- {
- char ifaceName[Smf::IF_NAME_MAX + 1];
- ifaceName[Smf::IF_NAME_MAX] = '\0';
- Smf::InterfaceGroup::Iterator ifacerator(*group);
- Smf::Interface* iface;
-
- // If we don't want to see PUSH groups, uncomment following two lines ...
- // if (Smf::PUSH == group->GetForwardingMode()) // I think we want to skip these ...
- // continue;
- ss << std::left << std::setw(21) << group->GetName();
- ss << std::setw(9) << (group->IsTemplateGroup() ? "Template" : "Regular") << " ";
- std::string relayType;
- switch (group->GetRelayType())
- {
- case Smf::INVALID: relayType="Invalid"; break;
- case Smf::CF: relayType="cf"; break;
- case Smf::S_MPR: relayType="s_mpr"; break;
- case Smf::E_CDS: relayType="e_cds"; break;
- case Smf::MPR_CDS: relayType="mpr_cds"; break;
- case Smf::NS_MPR: relayType="ns_mpr"; break;
- }
- ss << std::setw(9) << relayType << " ";
- switch (group->GetForwardingMode())
- {
- case Smf::PUSH: relayType="Push"; break;
- case Smf::MERGE: relayType="Merge"; break;
- case Smf::RELAY: relayType="Relay"; break;
- }
- ss << std::setw(14) << relayType << " ";
- bool firstInterface = true;
- while (NULL != (iface = ifacerator.GetNextInterface()))
- {
- ProtoNet::GetInterfaceName(iface->GetIndex(), ifaceName, Smf::IF_NAME_MAX);
- ss << ( firstInterface ? "" : ",") << ifaceName;
- firstInterface = false;
- }
- if (group->GetElasticMulticast())
- ss << ", Elastic";
- if (group->GetAdaptiveRouting())
- ss << ", Adaptive";
- ss << "\n";
- }
- ss << "\n";
- if (server_pipe.IsOpen())
- {
- unsigned int numBytes = ss.str().size();
- if (!server_pipe.Send(ss.str().c_str(), numBytes))
- {
- PLOG(PL_ERROR, "SmfApp::OnCommand(info) error sending info to smf server\n");
- return;
- }
- }
- else
- {
- fprintf(stderr, "Server pipe is NOT open\n");
- PLOG(PL_WARN, "SmfApp::OnCommand(info) warning: unable to connect to smfServer\n");
- }
+ ReplyInfo(false);
}
- else if (!strncmp("jsonStats", cmd, len)) // just checking stats ...
+ else if (!strncmp("jsonStats", cmd, len))
{
- std::ostringstream ss;
- ss << "[";
- if (server_pipe.IsOpen())
- {
- Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
- Smf::Interface* nextIface;
- bool comma = false;
- while (NULL != (nextIface = iterator.GetNextItem()))
- {
- ss << (comma ? "," : "") << "{";
- ss << "\"interface\":\"" << nextIface->GetNameStr() << "\",";
- ss << "\"flows\":\"" << nextIface->GetFlowCount() << "\",";
- ss << "\"recv\":\"" << nextIface->GetRecvCount() << "\",";
- ss << "\"mrcv\":\"" << nextIface->GetMcastCount() << "\",";
- ss << "\"sent\":\"" << nextIface->GetSentCount() << "\",";
- ss << "\"retr\":\"" << nextIface->GetRetransmissionCount() << "\",";
- ss << "\"fwd\":\"" << nextIface->GetForwardCount() << "\",";
- ss << "\"dups\":\"" << nextIface->GetDuplicateCount() << "\",";
- ss << "\"asym\":\"" << nextIface->GetAsymCount() << "\",";
- ss << "\"queue\":\"" << nextIface->GetQueueLength() << "\"";
- ss << "}";
- comma = true;
- }
- ss << "]\n";
- unsigned int numBytes = ss.str().size();
- if (!server_pipe.Send(ss.str().c_str(), numBytes))
- {
- PLOG(PL_ERROR, "SmfApp::OnCommand(jsonStats) error sending jsonStats to smf server\n");
- return;
- }
- }
- else
- {
- fprintf(stderr, "Server pipe is not open for stats\n");
- return;
- }
+ ReplyStats(true);
}
- else if (!strncmp("interfaces", cmd, len)) // checking interfaces
+ else if (0 == strcmp(cmd, "interfacesj"))
{
- std::ostringstream ss;
- if (server_pipe.IsOpen())
- {
- Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
- Smf::Interface* nextIface;
- ss << "Flags: L = Layered, T = Tunnel, I = IGMP Proxy, S = Shadowing\n\n";
- ss << "Interface Fwd Method Flags\n";
- ss << "---------------- ---------- -----\n";
- while (NULL != (nextIface = iterator.GetNextItem()))
- {
- ss << std::left << std::setw(16) << nextIface->GetNameStr() << " ";
- ss << std::setw(12);
-#ifdef ELASTIC_MCAST
- if (nextIface->GetElasticMulticast()) {
-
- if (mcast_controller.GetDefaultForwardingStatus() == MulticastFIB::HYBRID)
- ss << "Advertise";
- else
- ss << "Elastic";
- } else ss << "Flood";
-#else
- ss << "Flood";
-#endif // ELASTIC_MCAST
-
- std::setw(1);
- if (nextIface->IsLayered()) ss << "L";
- if (nextIface->IsTunnel()) ss << "T";
- if (nextIface->IsIgmpProxy()) ss << "I";
- InterfaceMechanism* mech = static_cast(nextIface->GetExtension());
- if ((NULL != mech) && mech->IsShadowing()) ss << "S";
- ss << "\n";
- }
- }
- unsigned int numBytes = ss.str().size();
- if (!server_pipe.Send(ss.str().c_str(), numBytes))
- {
- PLOG(PL_ERROR, "SmfApp::OnCommand(interfaces) error sending interfaces to smf server\n");
- return;
- }
+ ReplyInterfaces(true);
}
- else if (!strncmp("interfacesj", cmd, len)) // checking interfaces
+ else if (!strncmp("interfaces", cmd, len))
{
- std::ostringstream ss;
- if (server_pipe.IsOpen())
- {
- Smf::InterfaceList::Iterator iterator(smf.AccessInterfaceList());
- Smf::Interface* nextIface;
- bool comma = false;
-
- ss << "[";
- while (NULL != (nextIface = iterator.GetNextItem()))
- {
- ss << (comma ? "," : "") << "{";
- ss << "\"Interface\" : \"" << nextIface->GetNameStr() << "\",";
- ss << "\"FwdMethod\" : \"";
-#ifdef ELASTIC_MCAST
- if (nextIface->GetElasticMulticast()) {
- if (mcast_controller.GetDefaultForwardingStatus() == MulticastFIB::HYBRID)
- ss << "Advertise";
- else
- ss << "Elastic";
- } else ss << "Flood";
-#else
- ss << "Flood";
-#endif // ELASTIC_MCAST
- ss << "\",";
-
- ss << "\"Flags\" : \"";
- if (nextIface->IsLayered()) ss << "L";
- if (nextIface->IsTunnel()) ss << "T";
- if (nextIface->IsIgmpProxy()) ss << "I";
- InterfaceMechanism* mech = static_cast(nextIface->GetExtension());
- if ((NULL != mech) && mech->IsShadowing()) ss << "S";
- ss << "\"}";
- comma = true;
- }
- ss << "]\n";
- }
- unsigned int numBytes = ss.str().size();
- if (!server_pipe.Send(ss.str().c_str(), numBytes))
- {
- PLOG(PL_ERROR, "SmfApp::OnCommand(interfaces) error sending interfaces to smf server\n");
- return;
- }
+ ReplyInterfaces(false);
}
#ifdef ELASTIC_MCAST
- else if (!strncmp("brfgroups", cmd, len) || !strncmp("brfgroupsj", cmd, len)) // checking groups brief
+ else if (!strncmp("brfgroups", cmd, len) || !strncmp("brfgroupsj", cmd, len))
{
- std::ostringstream ss;
- bool useJson = cmd[len-1] == 'j';
-
- mcast_controller.DumpGroups(true, useJson, ss);
- unsigned int numBytes = ss.str().size();
- if (!server_pipe.Send(ss.str().c_str(), numBytes))
- {
- PLOG(PL_ERROR, "SmfApp::OnCommand(brfgroups) error sending brfgroups to smf server\n");
- return;
- }
+ ReplyGroups(cmd[len-1] == 'j', true);
}
- else if (!strncmp("groups", cmd, len) || !strncmp("groupsj", cmd, len)) // checking stats groups
+ else if (!strncmp("groups", cmd, len) || !strncmp("groupsj", cmd, len))
{
- std::ostringstream ss;
- bool useJson = cmd[len-1] == 'j';
-
- mcast_controller.DumpGroups(false, useJson, ss);
- unsigned int numBytes = ss.str().size();
- if (!server_pipe.Send(ss.str().c_str(), numBytes))
- {
- PLOG(PL_ERROR, "SmfApp::OnCommand(groups) error sending groups to smf server\n");
- return;
- }
+ ReplyGroups(cmd[len-1] == 'j', false);
}
#endif // ELASTIC_MCAST
@@ -6921,8 +7931,8 @@ void SmfApp::OnPktCapture(ProtoChannel& theChannel,
if (ProtoNet::IFACE_GRE == cap.GetInterfaceType())
{
// Create placeholder Ethernet header for packet received via GRE tunnel
- // Note the src/dst MAC addresses will be null for now
- // Note HandleInboundPacket() will populate Ethernet src/dst header fields later as needed
+ // src/dst MAC are null here; HandleInboundPacket() sets the dest
+ // for overlay multicast from the inner IP destination.
ProtoPktETH ethPkt;
ethPkt.InitIntoBuffer(ethBuffer, 14);
ethPkt.SetType(ProtoPktETH::IP);
@@ -7051,6 +8061,27 @@ bool SmfApp::SendFrame(unsigned int ifaceIndex, char* frameBuffer, unsigned int
return SendFrame(*iface, frameBuffer, frameLength);
} // end SmfApp::SendFrame()
+bool SmfApp::SendFrameTo(unsigned int ifaceIndex, char* frameBuffer, unsigned int frameLength,
+ const ProtoAddress& dest)
+{
+ Smf::Interface* iface = smf.GetInterface(ifaceIndex);
+ if (NULL == iface)
+ return false;
+ if (!dest.IsValid() || dest.HostIsEqual(PROTO_ADDR_ANY) || dest.HostIsEqual(PROTO_ADDR_ANY6) ||
+ (ProtoAddress::ETH == dest.GetType()))
+ return SendFrame(*iface, frameBuffer, frameLength);
+
+ InterfaceMechanism* mech = static_cast(iface->GetExtension());
+ if (NULL == mech)
+ return SendFrame(*iface, frameBuffer, frameLength);
+ CidElement* elem = mech->GetPrincipalElement();
+ if ((NULL == elem) || (ProtoNet::IFACE_GRE != elem->GetProtoCap().GetInterfaceType()))
+ return SendFrame(*iface, frameBuffer, frameLength);
+
+ unsigned int numBytes = frameLength;
+ return mech->SendGreToRemote(elem->GetProtoCap(), frameBuffer, frameLength, dest, numBytes);
+} // end SmfApp::SendFrameTo()
+
// Forward IP packet encapsulated in ETH frame using "ProtoCap" (i.e. pcap or similar) device
bool SmfApp::SendFrame(Smf::Interface& iface, char* frameBuffer, unsigned int frameLength)
{
@@ -7254,8 +8285,15 @@ bool SmfApp::HandleInboundPacket(UINT32* alignedBuffer, unsigned int numBytes, P
bool srcCapIsGRE = (ProtoNet::IFACE_GRE == srcCap.GetInterfaceType());
if (srcCapIsGRE)
{
- // This will be IP instead of ETH and may be INADDR_ANY for mGRE tunnels
+ // Configured remote is INADDR_ANY for mGRE; use the per-packet
+ // GRE outer source so each neighbor has a distinct previous hop.
prevHopAddr = srcCap.GetTunnelRemoteAddr();
+ if (TunnelAddrUnspecified(prevHopAddr))
+ {
+ const ProtoAddress& pktRemote = srcCap.GetPacketRemoteAddr();
+ if (pktRemote.IsValid() && !TunnelAddrUnspecified(pktRemote))
+ prevHopAddr = pktRemote;
+ }
nextHopAddr = srcCap.GetTunnelLocalAddr();
}
else
@@ -7319,6 +8357,17 @@ bool SmfApp::HandleInboundPacket(UINT32* alignedBuffer, unsigned int numBytes, P
}
if (dstAddr.IsUnicast()) isUnicast = true;
+ // GRE receive builds a placeholder Ethernet header with a null dest.
+ // Forwarding onto a real LAN needs the IP-mapped multicast MAC (01:00:5e:... /
+ // 33:33:...), not 00:00:00:00:00:00. ProtoCap::Forward() only rewrites src.
+ if (srcCapIsGRE && !isUnicast)
+ {
+ ProtoAddress ethDst;
+ ethDst.GetEthernetMulticastAddress(dstAddr);
+ if (ethDst.IsValid())
+ ethPkt.SetDstAddr(ethDst);
+ }
+
// Some IGMP snooping test code (TBD - handle IPv6 too)
bool igmpSnoop = false;
if (igmpSnoop && (ProtoPktIP::IGMP == protocol))
@@ -7397,12 +8446,7 @@ bool SmfApp::HandleInboundPacket(UINT32* alignedBuffer, unsigned int numBytes, P
// doesn't seem to be necessary to fix
ethPkt.SetDstAddr(vif->GetHardwareAddress());
}
- else
- {
- // Fix ETH dstMacAddr for multicast
- dstMacAddr.GetEthernetMulticastAddress(dstAddr);
- ethPkt.SetDstAddr(dstMacAddr);
- }
+ // else: multicast dest already set from the GRE overlay IP dest above
}
else
{
@@ -7558,6 +8602,19 @@ void SmfApp::MonitorEventHandler(ProtoChannel& theChannel,
unsigned int ifIndex = theEvent.GetInterfaceIndex();
const char* ifName = theEvent.GetInterfaceName();
+ if ((ProtoNet::Monitor::Event::IFACE_NEIGH_NEW == theEvent.GetType()) ||
+ (ProtoNet::Monitor::Event::IFACE_NEIGH_DELETE == theEvent.GetType()))
+ {
+ Smf::Interface* neighIface = smf.GetInterface(ifIndex);
+ if ((NULL != neighIface) && neighIface->GetTunnelLearnDynamic())
+ {
+ UpdateLearnedTunnelPeer(*neighIface, theEvent.GetAddress(),
+ theEvent.GetAuxAddress(), theEvent.GetFlags(),
+ ProtoNet::Monitor::Event::IFACE_NEIGH_DELETE == theEvent.GetType());
+ }
+ continue;
+ }
+
// Is this an interface we care about?
// a) Is it one of our interfaces?
Smf::Interface* iface = smf.GetInterface(ifIndex);
@@ -7640,7 +8697,7 @@ void SmfApp::MonitorEventHandler(ProtoChannel& theChannel,
iface->SetTunnelLocalAddress(localAddr);
iface->SetTunnelRemoteAddress(remoteAddr);
// not "remoteAddr" may be INADDR_ANY for mGRE tunnels
- smf.AddTunnelInfo(ifIndex, localAddr, remoteAddr);
+ smf.AddTunnelInfo(ifIndex, localAddr, remoteAddr, false);
}
else
{
diff --git a/src/common/smf.cpp b/src/common/smf.cpp
index 8d8a865..b20f8a3 100644
--- a/src/common/smf.cpp
+++ b/src/common/smf.cpp
@@ -79,6 +79,7 @@ Smf::Interface::Interface(unsigned int ifIndex, const char *ifName)
sent_count(0), retr_count(0), recv_count(0),
mrcv_count(0), dups_count(0), asym_count(0), fwd_count(0), extension(NULL)
{
+ tunnel_learn_dynamic = false;
}
Smf::Interface::~Interface()
@@ -131,9 +132,23 @@ void Smf::Interface::Destroy()
assoc_source_list.Destroy(); // this deletes the items which also removes them from the sources' target lists
// Destroy our target list
assoc_target_list.Destroy();
+ ClearLearnedOverlays();
} // end Smf::Interface::Destroy()
+void Smf::Interface::ClearLearnedOverlays()
+{
+ ProtoAddress overlay;
+ ProtoAddressList::Iterator it(learned_overlays);
+ while (it.GetNextAddress(overlay))
+ {
+ const ProtoAddress* underlay =
+ static_cast(learned_overlays.GetUserData(overlay));
+ delete const_cast(underlay);
+ }
+ learned_overlays.Destroy();
+} // end Smf::Interface::ClearLearnedOverlays()
+
bool Smf::Interface::AddAssociate(InterfaceGroup& ifaceGroup, Interface& iface)
{
// (TBD) Should we verify that there isn't already an "Associate"
@@ -1146,24 +1161,41 @@ int Smf::ProcessPacket(ProtoPktIP& ipPkt, // input/output - the
PLOG(PL_DETAIL, "Smf::ProcessPacket(): IPv4 Packet detected: Length = %d.\n" , (UINT16)ipv4Pkt.GetLength());
PLOG(PL_DETAIL, "Smf::ProcessPacket(): IPv4 Packet detected: FragmentOffset = %d.\n" , (UINT16)ipv4Pkt.GetFragmentOffset());
#ifdef ELASTIC_MCAST
-
+ // EM control may be unicast on mGRE (inner dest is the peer overlay).
+ bool elasticCtl = false;
+ if (ProtoPktIP::UDP == ipv4Pkt.GetProtocol())
+ {
+ ProtoPktUDP peekUdp;
+ if (peekUdp.InitFromPacket(ipv4Pkt) &&
+ (peekUdp.GetDstPort() == ElasticAck::ELASTIC_PORT))
+ elasticCtl = true;
+ }
#endif // ELASTIC_MCAST
- if (!dstIp.IsMulticast() && !GetUnicastEnabled() && !GetAdaptiveRouting()) // only forward multicast dst, unless unicast enabled
+ if (!dstIp.IsMulticast() && !GetUnicastEnabled() && !GetAdaptiveRouting()
+#ifdef ELASTIC_MCAST
+ && !elasticCtl
+#endif // ELASTIC_MCAST
+ ) // only forward multicast dst, unless unicast enabled
{
PLOG(PL_DETAIL, "Smf::ProcessPacket() skipping non-multicast IPv4 pkt\n");
return 0;
}
+#ifdef ELASTIC_MCAST
+ else if (elasticCtl ||
+ (dstIp.IsLinkLocal() &&
+ (!is_tunnel || dstIp.HostIsEqual(ElasticAck::ELASTIC_ADDR))))
+#else
else if (dstIp.IsLinkLocal() && !is_tunnel) // don't forward if link-local dst
+#endif // ELASTIC_MCAST
{
#ifdef ELASTIC_MCAST
// TBD - use non-link local address for ElasticMcast control messages so that ACK/NACK
// message to enable assymmetric/non-reciprocal link topology support
// Is this an ElasticMulticast ACK? (if so, notify controller)
- if (dstIp.HostIsEqual(ElasticAck::ELASTIC_ADDR) &&
- (ProtoPktIP::UDP == ipv4Pkt.GetProtocol()))
+ if (ProtoPktIP::UDP == ipv4Pkt.GetProtocol())
{
ProtoPktUDP udpPkt;
if (udpPkt.InitFromPacket(ipv4Pkt) && (udpPkt.GetDstPort() == ElasticAck::ELASTIC_PORT))
@@ -1235,6 +1267,13 @@ int Smf::ProcessPacket(ProtoPktIP& ipPkt, // input/output - the
PLOG(PL_DEBUG, "Smf::ProcessPacket() EM_ACK mapped to GRE interface '%s'\n",
upstreamIface->GetNameStr());
}
+ else if ((0 != (upstreamIndex = FindInterfaceByLocalEndpoint(upstreamAddr))) &&
+ (NULL != (upstreamIface = GetInterface(upstreamIndex))))
+ {
+ // 2c) mGRE: ACK upstream is our tunnel local (or overlay IP)
+ PLOG(PL_DEBUG, "Smf::ProcessPacket() EM_ACK mapped to local endpoint interface '%s'\n",
+ upstreamIface->GetNameStr());
+ }
else if ((0 != (upstreamIndex = GetInterfaceIndex(upstreamAddr))) &&
(NULL != (upstreamIface = GetInterface(upstreamIndex))))
{
@@ -1253,6 +1292,11 @@ int Smf::ProcessPacket(ProtoPktIP& ipPkt, // input/output - the
// else // else not for me
}
} // end for (UINT8 i = 0; i < upstreamCount
+ // mGRE unicast ACK: arrived on this tunnel and mapping
+ // missed (no local endpoint match). The GRE dest already
+ // selected this node, so apply the ACK to srcIface.
+ if ((NULL == upstreamIface) && srcIface.IsGRE())
+ mcast_controller->HandleAck(elasticAck, &srcIface, srcIp, prevHopAddr);
break;
}
case ElasticMsg::ADV:
@@ -2030,6 +2074,8 @@ int Smf::ProcessPacket(ProtoPktIP& ipPkt, // input/output - the
// If there is an ElasticMcast interface group, this will
// be looked up (or created as needed for new flows)
MulticastFIB::Entry* fibEntry = NULL;
+ bool recvNoted = false;
+ MulticastFIB::TokenBucket* bucket = NULL;
#endif // ELASTIC_MCAST
@@ -2333,6 +2379,7 @@ int Smf::ProcessPacket(ProtoPktIP& ipPkt, // input/output - the
#endif // ADAPTIVE_ROUTING
#ifdef ELASTIC_MCAST
+ bucket = NULL;
if (!dstIp.IsMulticast() && IsOwnAddress(dstIp))
{
// Don't forward unicast packets destined to self
@@ -2368,8 +2415,14 @@ int Smf::ProcessPacket(ProtoPktIP& ipPkt, // input/output - the
}
}
} // end if (NULL == fibEntry)
+ if (!recvNoted && dstIp.IsMulticast() &&
+ !dstIp.HostIsEqual(ElasticAck::ELASTIC_ADDR))
+ {
+ fibEntry->NoteRecv(currentTick);
+ recvNoted = true;
+ }
// Get (or create if needed) the token bucket for this outbound iface
- MulticastFIB::TokenBucket* bucket = fibEntry->GetBucket(dstIface.GetIndex());
+ bucket = fibEntry->GetBucket(dstIface.GetIndex());
if (NULL != bucket)
{
// Check if the flow passes the bucket's rate limit test
@@ -2398,8 +2451,10 @@ int Smf::ProcessPacket(ProtoPktIP& ipPkt, // input/output - the
if (((ttl > 1) || is_tunnel || outbound) && ((unsigned int)dstCount < dstIfArraySize))
{
dstIfArray[dstCount++] = dstIface.GetIndex();
-
-
+#ifdef ELASTIC_MCAST
+ if (NULL != bucket)
+ bucket->NoteSent(currentTick);
+#endif // ELASTIC_MCAST
}
PLOG(PL_DETAIL, "Smf::ProcessPacket(): Preparing to forward! DstCount = %d \n", dstCount );
}
@@ -2792,6 +2847,27 @@ MulticastFIB::Entry* Smf::UpdateElasticRouting(unsigned int cu
}
}
mcast_fib.InsertEntry(*fibEntry);
+ // IGMP/static memberships are (*,G). Seed last-hop FORWARD on
+ // this new (S,G) so the first packet is not stuck at LIMIT.
+ if (NULL != mcast_controller)
+ {
+ ProtoAddress dstIp;
+ flowDescription.GetDstAddr(dstIp);
+ ProtoFlow::Description dstOnly(dstIp);
+ MulticastFIB::MembershipTable::Iterator memIt(
+ mcast_controller->AccessMembershipTable(), &dstOnly,
+ ProtoFlow::Description::FLAG_DST);
+ MulticastFIB::Membership* membership;
+ while (NULL != (membership = memIt.GetNextEntry()))
+ {
+ if (membership->FlagIsSet(MulticastFIB::Membership::MANAGED) ||
+ membership->FlagIsSet(MulticastFIB::Membership::STATIC))
+ {
+ fibEntry->SetForwardingStatus(membership->GetInterfaceIndex(),
+ MulticastFIB::FORWARD, true);
+ }
+ }
+ }
// Put the new, dynamically detected flow in our "active_list"
mcast_fib.ActivateFlow(*fibEntry, currentTick);
updateController = true;
@@ -3223,6 +3299,118 @@ unsigned int Smf::UpdateUpstreamHistory(unsigned int currentTi
}
} // end Smf::SendNack()
+bool Smf::FindTunnelUnicastPeer(unsigned int ifaceIndex, const ProtoAddress& addr,
+ ProtoAddress& dest)
+{
+ if (!addr.IsValid() || (ProtoAddress::ETH == addr.GetType()) ||
+ addr.IsMulticast() ||
+ addr.HostIsEqual(PROTO_ADDR_ANY) || addr.HostIsEqual(PROTO_ADDR_ANY6))
+ return false;
+ InterfaceInfoTable::Iterator iterator(iface_info_table);
+ InterfaceInfo* info;
+ while (NULL != (info = iterator.GetNextItem()))
+ {
+ if (info->GetIndex() != ifaceIndex)
+ continue;
+ const ProtoAddress& remote = info->GetRemoteAddress();
+ if (remote.IsValid() && remote.IsUnicast() && remote.HostIsEqual(addr))
+ {
+ dest = remote;
+ return true;
+ }
+ }
+ Interface* iface = GetInterface(ifaceIndex);
+ if (NULL != iface)
+ {
+ const ProtoAddress* underlay =
+ static_cast(iface->AccessLearnedOverlays().GetUserData(addr));
+ if ((NULL != underlay) && underlay->IsValid() && underlay->IsUnicast())
+ {
+ dest = *underlay;
+ return true;
+ }
+ }
+ return false;
+}
+
+namespace {
+struct OverlayNeighCtx
+{
+ const ProtoAddress* underlay;
+ ProtoAddress* overlay;
+ bool found;
+};
+
+bool OverlayNeighHandler(unsigned int /*ifIndex*/,
+ const ProtoAddress& dst,
+ const ProtoAddress& lladdr,
+ unsigned short /*ndmState*/,
+ void* userData)
+{
+ OverlayNeighCtx* ctx = static_cast(userData);
+ if ((NULL == ctx) || ctx->found)
+ return false;
+ if (lladdr.IsValid() && lladdr.IsUnicast() &&
+ lladdr.HostIsEqual(*ctx->underlay) &&
+ dst.IsValid() && dst.IsUnicast() &&
+ (ProtoAddress::ETH != dst.GetType()))
+ {
+ *ctx->overlay = dst;
+ ctx->found = true;
+ return false;
+ }
+ return true;
+}
+} // namespace
+
+bool Smf::FindOverlayForUnderlay(unsigned int ifaceIndex, const ProtoAddress& underlay,
+ ProtoAddress& overlay)
+{
+ Interface* iface = GetInterface(ifaceIndex);
+ if ((NULL == iface) || !underlay.IsValid() || !underlay.IsUnicast())
+ return false;
+ ProtoAddressList& learned = iface->AccessLearnedOverlays();
+ ProtoAddress ov;
+ ProtoAddressList::Iterator it(learned);
+ while (it.GetNextAddress(ov))
+ {
+ const ProtoAddress* ul = static_cast(learned.GetUserData(ov));
+ if ((NULL != ul) && ul->HostIsEqual(underlay))
+ {
+ overlay = ov;
+ return true;
+ }
+ }
+ // Explicit-map receivers do not populate learned_overlays. The
+ // static NBMA table is kernel ip neigh on gre1 (overlay -> underlay).
+ OverlayNeighCtx ctx;
+ ctx.underlay = &underlay;
+ ctx.overlay = &overlay;
+ ctx.found = false;
+ ProtoNet::GetInterfaceNeighbors(ifaceIndex, OverlayNeighHandler, &ctx);
+ return ctx.found;
+}
+
+unsigned int Smf::FindInterfaceByLocalEndpoint(const ProtoAddress& addr)
+{
+ if (!addr.IsValid() ||
+ addr.HostIsEqual(PROTO_ADDR_ANY) ||
+ addr.HostIsEqual(PROTO_ADDR_ANY6))
+ return 0;
+ InterfaceList::Iterator it(iface_list);
+ Interface* iface;
+ while (NULL != (iface = it.GetNextInterface()))
+ {
+ const ProtoAddress& tunLocal = iface->GetTunnelLocalAddress();
+ if (tunLocal.IsValid() && tunLocal.HostIsEqual(addr))
+ return iface->GetIndex();
+ const ProtoAddress& ip = iface->GetIpAddress();
+ if (ip.IsValid() && (ProtoAddress::ETH != ip.GetType()) && ip.HostIsEqual(addr))
+ return iface->GetIndex();
+ }
+ return 0;
+}
+
bool Smf::SendAck(unsigned int ifaceIndex, // interface it goes out on
const ProtoAddress& upstreamAddr, // upstream to address it to
const ProtoFlow::Description& flowDescription)
@@ -3243,16 +3431,41 @@ bool Smf::SendAck(Interface& iface, // interface it go
// Buid Elastic Ack message (IPv4 only at moment)
const ProtoAddress& dstMac = (ProtoAddress::ETH == upstreamAddr.GetType()) ? upstreamAddr : ElasticNack::ELASTIC_MAC;
- if (iface.GetIpAddress().GetType() == ProtoAddress::INVALID)
+ const bool mgre = iface.IsGRE() &&
+ (!iface.GetTunnelRemoteAddress().IsValid() ||
+ iface.GetTunnelRemoteAddress().HostIsEqual(PROTO_ADDR_ANY) ||
+ iface.GetTunnelRemoteAddress().HostIsEqual(PROTO_ADDR_ANY6));
+ // P2P GRE uses tunnel local as ACK src. mGRE/ETH use the iface IP
+ // (overlay on gre1); fall back to tunnel local if that is unset.
+ ProtoAddress srcIp;
+ if (iface.IsGRE() && !mgre)
+ srcIp = iface.GetTunnelLocalAddress();
+ else if (iface.GetIpAddress().IsValid() && (ProtoAddress::ETH != iface.GetIpAddress().GetType()))
+ srcIp = iface.GetIpAddress();
+ else
+ srcIp = iface.GetTunnelLocalAddress();
+ if (!srcIp.IsValid() || (ProtoAddress::ETH == srcIp.GetType()))
{
PLOG(PL_WARN, "Smf::SendAck() no IP address on interface %s!\n", iface.GetNameStr());
return false;
}
- // The EM_ACK srcIp depends on whether interface is ETH, GRE, or mGRE
- const ProtoAddress& srcIp = (iface.IsGRE() &&
- !iface.GetTunnelRemoteAddress().HostIsEqual(PROTO_ADDR_ANY) &&
- !iface.GetTunnelRemoteAddress().HostIsEqual(PROTO_ADDR_ANY6)) ?
- iface.GetTunnelLocalAddress() : iface.GetIpAddress();
+ // Ethernet keeps 224.0.0.55. mGRE unicasts the inner dest to the
+ // peer overlay so gre1 delivers it (224.0.0.55 and the flow src
+ // LAN address are not local on the tunnel).
+ ProtoAddress ackDst = ElasticAck::ELASTIC_ADDR;
+ ProtoAddress mgrePeer;
+ if (mgre)
+ {
+ ProtoAddress overlay;
+ if (FindTunnelUnicastPeer(iface.GetIndex(), upstreamAddr, mgrePeer))
+ ;
+ else if (upstreamAddr.IsValid() && upstreamAddr.IsUnicast() &&
+ (ProtoAddress::ETH != upstreamAddr.GetType()))
+ mgrePeer = upstreamAddr;
+ if (mgrePeer.IsValid() &&
+ FindOverlayForUnderlay(iface.GetIndex(), mgrePeer, overlay))
+ ackDst = overlay;
+ }
UINT32 buffer[1416/4];
unsigned int bufferLen = 1416;
unsigned int frameMax = bufferLen - 2; // offset by 2 bytes to maintain alignment for ProtoPktIP
@@ -3266,7 +3479,7 @@ bool Smf::SendAck(Interface& iface, // interface it go
ip4Pkt.SetTTL(5);
ip4Pkt.SetProtocol(ProtoPktIP::UDP);
ip4Pkt.SetSrcAddr(srcIp);
- ip4Pkt.SetDstAddr(ElasticAck::ELASTIC_ADDR);
+ ip4Pkt.SetDstAddr(ackDst);
ProtoPktUDP udpPkt(ip4Pkt.AccessPayload(), ip4Pkt.GetBufferLength() - ip4Pkt.GetHeaderLength() - ProtoPktUMP::GetOptionLength(), false);
udpPkt.SetSrcPort(ElasticAck::ELASTIC_PORT);
udpPkt.SetDstPort(ElasticAck::ELASTIC_PORT);
@@ -3327,11 +3540,24 @@ bool Smf::SendAck(Interface& iface, // interface it go
{
PLOG(PL_DEBUG, "nrlsmf: sending EM_ACK (len:%u) for flow \"", ethPkt.GetLength());
flowDescription.Print(); // to debug output or log
- PLOG(PL_ALWAYS, " to relay %s via interface index %d\n", upstreamAddr.GetHostString(), iface.GetIndex());
+ PLOG(PL_ALWAYS, " to relay %s via interface index %d dest %s\n",
+ upstreamAddr.GetHostString(), iface.GetIndex(), ackDst.GetHostString());
}
// TBD - Implement ACK rate limiter by bundling multiple flow acks for common upstream relay
// (i.e. do this with a timer and some sort of helper classes)
+ // mGRE ACK dest mirrors data inject: unicast if that peer is in the
+ // map (or learned), otherwise the underlay multicast group.
+ if (mgre)
+ {
+ ProtoAddress mcastUnderlay;
+ if (mgrePeer.IsValid())
+ return output_mechanism->SendFrameTo(iface.GetIndex(), (char*)ethPkt.GetBuffer(),
+ ethPkt.GetLength(), mgrePeer);
+ if (GetTunnelMulticastRemote(iface.GetIndex(), mcastUnderlay))
+ return output_mechanism->SendFrameTo(iface.GetIndex(), (char*)ethPkt.GetBuffer(),
+ ethPkt.GetLength(), mcastUnderlay);
+ }
return output_mechanism->SendFrame(iface.GetIndex(), (char*)ethPkt.GetBuffer(), ethPkt.GetLength());
} // end Smf::SendAck()
diff --git a/src/common/smfIgmp.cpp b/src/common/smfIgmp.cpp
index 1909f62..2b6195e 100644
--- a/src/common/smfIgmp.cpp
+++ b/src/common/smfIgmp.cpp
@@ -37,13 +37,7 @@ bool SmfIgmp::Open(bool withFRR)
wpipe = p[1]; // For writing IGMP updates locally
if (withFRR)
- { // Using FRR, so active the timer to check FRR
- DoUpdate(update_timer);
- if (!update_timer.IsActive())
- {
- timer_mgr.ActivateTimer(update_timer);
- }
- }
+ EnableFrrPolling();
else if (update_timer.IsActive())
{ // Not using FRR, so make sure the timer is inactive
update_timer.Deactivate();
@@ -59,6 +53,13 @@ bool SmfIgmp::Open(bool withFRR)
return true;
}
+void SmfIgmp::EnableFrrPolling()
+{
+ DoUpdate(update_timer);
+ if (!update_timer.IsActive())
+ timer_mgr.ActivateTimer(update_timer);
+}
+
void SmfIgmp::Close()
{
// This must be called first
diff --git a/tests/mutests/1hop_smf/README.md b/tests/mutests/1hop_smf/README.md
new file mode 100644
index 0000000..a0824f7
--- /dev/null
+++ b/tests/mutests/1hop_smf/README.md
@@ -0,0 +1,38 @@
+# Single-router nrlsmf mutests
+
+One router (r0) bridging two host LANs, exercising nrlsmf's basic
+forwarding modes and CLI.
+
+```
+h0 -- lan0 --\
+ r0
+h1 -- lan1 --/
+```
+
+## Tests
+
+| File | What it covers |
+|------|-----------------|
+| `mutest_smf_cli.py` | nrlsmf command-line parsing sanity checks (`help`, `version`, invalid/ambiguous commands) — no host traffic |
+| `mutest_smf_merge.py` | `merge eth0,eth1` — forced two-interface relay (Gateway Command) |
+| `mutest_smf_cf.py` | `add net,cf,eth0,eth1` — classical flooding, including the implicit `push:eth0`/`push:eth1` sub-groups it creates |
+| `mutest_smf_elastic.py` | Elastic Multicast (EM) overlaid on the same `cf` group — data-plane rate limiting instead of blind flooding |
+| `mutest_smf_advertise.py` | EM `advertise` mode — confirms EM_ADV control messages are actually sent (224.0.0.55:5555) |
+
+Each file's docstring explains what its mode means and how it differs
+from the others, and each is self-contained (starts and stops its own
+nrlsmf instance and any iperf/tcpdump processes it needs).
+
+## Run
+
+From `tests/mutests` (requires root, FRR, `nrlsmf` on PATH):
+
+```bash
+sudo mutest 1hop_smf
+# or one file:
+sudo mutest 1hop_smf/mutest_smf_cli.py
+sudo mutest 1hop_smf/mutest_smf_merge.py
+sudo mutest 1hop_smf/mutest_smf_cf.py
+sudo mutest 1hop_smf/mutest_smf_elastic.py
+sudo mutest 1hop_smf/mutest_smf_advertise.py
+```
diff --git a/tests/mutests/1hop_smf/h2/etc.frr/daemons b/tests/mutests/1hop_smf/h0/etc.frr/daemons
similarity index 100%
rename from tests/mutests/1hop_smf/h2/etc.frr/daemons
rename to tests/mutests/1hop_smf/h0/etc.frr/daemons
diff --git a/tests/mutests/1hop_smf/h0/etc.frr/frr.conf b/tests/mutests/1hop_smf/h0/etc.frr/frr.conf
new file mode 100644
index 0000000..8cee8d5
--- /dev/null
+++ b/tests/mutests/1hop_smf/h0/etc.frr/frr.conf
@@ -0,0 +1,7 @@
+log file /var/log/frr/frr.log
+!
+ip route 0.0.0.0/0 10.0.0.1
+!
+interface eth0
+ ip address 10.0.0.2/24
+!
diff --git a/tests/mutests/1hop_smf/h2/etc.frr/vtysh.conf b/tests/mutests/1hop_smf/h0/etc.frr/vtysh.conf
similarity index 100%
rename from tests/mutests/1hop_smf/h2/etc.frr/vtysh.conf
rename to tests/mutests/1hop_smf/h0/etc.frr/vtysh.conf
diff --git a/tests/mutests/1hop_smf/h1/etc.frr/frr.conf b/tests/mutests/1hop_smf/h1/etc.frr/frr.conf
index 7450fea..3283210 100644
--- a/tests/mutests/1hop_smf/h1/etc.frr/frr.conf
+++ b/tests/mutests/1hop_smf/h1/etc.frr/frr.conf
@@ -5,6 +5,3 @@ ip route 0.0.0.0/0 10.0.1.1
interface eth0
ip address 10.0.1.2/24
!
-
-
-
diff --git a/tests/mutests/1hop_smf/h2/etc.frr/frr.conf b/tests/mutests/1hop_smf/h2/etc.frr/frr.conf
deleted file mode 100644
index 3602427..0000000
--- a/tests/mutests/1hop_smf/h2/etc.frr/frr.conf
+++ /dev/null
@@ -1,7 +0,0 @@
-log file /var/log/frr/frr.log
-!
-ip route 0.0.0.0/0 10.0.2.1
-!
-interface eth0
- ip address 10.0.2.2/24
-!
diff --git a/tests/mutests/1hop_smf/munet.yaml b/tests/mutests/1hop_smf/munet.yaml
index b2f943c..78db1bf 100644
--- a/tests/mutests/1hop_smf/munet.yaml
+++ b/tests/mutests/1hop_smf/munet.yaml
@@ -1,20 +1,27 @@
+# Single router bridging two host LANs.
+#
+# h0 -- lan0 --\
+# r0
+# h1 -- lan1 --/
+#
+# r0 is the router under test: all the nrlsmf forwarding-mode scenarios
+# in this directory run on r0, relaying multicast traffic between h0
+# and h1 in different ways. h0 and h1 are plain hosts, one on each side.
topology:
- networks:
- - name: hnet1
- - name: hnet2
-# simple network with two hosts and one router
-# h1 <-> r1 <-> h2
- nodes:
- - name: h1
- kind: frr
- connections:
- - to: hnet1
- - name: h2
- kind: frr
- connections:
- - to: hnet2
- - name: r1
- kind: frr
- connections:
- - to: hnet1
- - to: hnet2
+ networks:
+ - name: lan0
+ - name: lan1
+ nodes:
+ - name: h0
+ kind: frr
+ connections:
+ - to: lan0
+ - name: h1
+ kind: frr
+ connections:
+ - to: lan1
+ - name: r0
+ kind: frr
+ connections:
+ - to: lan0
+ - to: lan1
diff --git a/tests/mutests/1hop_smf/mutest_1hop_smf.py b/tests/mutests/1hop_smf/mutest_1hop_smf.py
deleted file mode 100644
index 64ae355..0000000
--- a/tests/mutests/1hop_smf/mutest_1hop_smf.py
+++ /dev/null
@@ -1,307 +0,0 @@
-"""Basic SMF mutest."""
-
-import subprocess
-
-from munet.mutest.userapi import match_step, step
-from munet.mutest.userapi import section
-from munet.mutest.userapi import test_step
-from munet.mutest.userapi import wait_step
-
-
-def pipe(cmd: str) -> subprocess.Popen:
- """Start a command and capture stdout/stderr through a pipe."""
- return subprocess.Popen(
- cmd,
- shell=True,
- stdout=subprocess.PIPE,
- stderr=subprocess.STDOUT,
- text=True,
- )
-
-
-def pipe_read(proc: subprocess.Popen, lines: int = 1) -> str:
- """Read up to 'lines' lines from a running process."""
- out = []
- if proc.stdout is None:
- return ""
- for _ in range(lines):
- line = proc.stdout.readline()
- if not line:
- break
- out.append(line)
- return "".join(out)
-
-
-def pipe_close(proc: subprocess.Popen, timeout: int = 5) -> None:
- """Terminate a process cleanly, then force kill if needed."""
- if proc.poll() is not None:
- return
- proc.terminate()
- try:
- proc.wait(timeout=timeout)
- except subprocess.TimeoutExpired:
- proc.kill()
- proc.wait()
-
-
-section("Verify interfaces are ready")
-
-for node in ("h1", "h2", "r1"):
- step(node, "ethtool -K eth0 rx off tx off")
-
-step("r1", "ethtool -K eth1 rx off tx off")
-
-wait_step(
- "h1",
- "ip -br addr show dev eth0 ",
- match="10.0.1.2/24",
- desc="IP address assigned to eth0",
-)
-wait_step(
- "h2",
- "ip -br addr show dev eth0 ",
- match="10.0.2.2/24",
- desc="IP address assigned to eth0",
-)
-
-wait_step(
- "r1",
- "ip -br addr show dev eth0 ",
- match="10.0.1.1/24",
- desc="IP address assigned to eth1",
-)
-
-wait_step(
- "r1",
- "ip -br addr show dev eth1 ",
- match="10.0.2.1/24",
- desc="IP address assigned to eth1",
-)
-
-section("Validate simple nrlsmf CLI arguments")
-
-help_output = step("r1", "sh -lc 'nrlsmf help; echo EXIT:$?'")
-test_step("Usage: nrlsmf" in help_output, "nrlsmf help prints usage", target="r1")
-test_step("EXIT:0" in help_output, "nrlsmf help exits successfully", target="r1")
-test_step("forward {on | off}" in help_output, "nrlsmf help lists forward option", target="r1")
-test_step("relay {on | off}" in help_output, "nrlsmf help lists relay option", target="r1")
-test_step("resequence {on | off}" in help_output, "nrlsmf help lists resequence option", target="r1")
-test_step("window {on | off}" in help_output, "nrlsmf help lists window option", target="r1")
-
-version_output = step("r1", "nrlsmf version")
-test_step(bool(version_output.strip()), "nrlsmf version prints non-empty output", target="r1")
-
-version_abbrev_output = step("r1", "sh -lc 'nrlsmf ver; echo EXIT:$?'")
-test_step("smf version:" in version_abbrev_output, "abbreviated 'ver' prints version", target="r1")
-test_step("EXIT:0" in version_abbrev_output, "abbreviated 'ver' exits successfully", target="r1")
-
-invalid_cmd_output = step("r1", "sh -lc 'nrlsmf nope; echo EXIT:$?'")
-test_step("Usage: nrlsmf" in invalid_cmd_output, "invalid command prints usage", target="r1")
-test_step("EXIT:0" not in invalid_cmd_output, "invalid command exits non-zero", target="r1")
-
-ambiguous_cmd_output = step("r1", "sh -lc 'nrlsmf r; echo EXIT:$?'")
-test_step("Usage: nrlsmf" in ambiguous_cmd_output, "ambiguous command prints usage", target="r1")
-test_step("EXIT:0" not in ambiguous_cmd_output, "ambiguous command exits non-zero", target="r1")
-
-# Merge flooding
-section("Start nrlsmf merge eth0,eth1 on r1 ")
-
-step("r1", "nrlsmf debug 4 merge eth0,eth1 &> nrlsmf-merge.log &")
-
-wait_step(
- "r1",
- 'pgrep -af "nrlsmf.*merge eth0,eth1"',
- match="merge eth0,eth1",
- desc="nrlsmf is started with merge eth0,eth1",
-)
-
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-merge.log',
- match='"merge" eth0,eth1',
- desc="nrlsmf-merge.log contains merge group for eth0,eth1",
-)
-
-wait_step(
- "r1",
- "nrlsmf --cli ping",
- match="pong",
- desc="nrlsmf --cli ping returns pong",
- timeout=10,
-)
-
-step(
- "h1",
- "iperf -u -T 4 -t 1000 -i 1 -b 8pps -l 1024 -e -c 239.0.0.1 &> iperf-client.log &",
-)
-
-wait_step(
- "h1",
- "tail -n1 iperf-client.log",
- match="8 pps",
- desc="Sending 239.0.0.1 at 8 pps",
-)
-
-step("h2", "iperf -u -T 4 -i 1 -s -e -B 239.0.0.1 > iperf-server.log 2>&1 &")
-
-wait_step(
- "h2",
- "tail -n1 iperf-server.log",
- match="8 pps",
- desc="Receiving 239.0.0.1 at full rate of 8 pps",
-)
-
-step("r1", "pkill nrlsmf")
-
-wait_step(
- "r1",
- 'pgrep -af "nrlsmf"',
- match="",
- desc="stopped nrlsmf",
-)
-
-# Classic flooding
-section("Start nrlsmf with classic flooding on r1 ")
-step(
- "r1",
- "nrlsmf debug 4 add net,cf,eth0,eth1 &> nrlsmf-cf.log &",
-)
-
-wait_step(
- "r1",
- 'pgrep -af "nrlsmf.*net,cf,eth0,eth1"',
- match="net,cf,eth0,eth1",
- desc="nrlsmf is started with classic flooding group",
-)
-
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-cf.log',
- match='"net" eth0,eth1',
- desc='nrlsmf-cf.log contains group "net" eth0,eth1',
-)
-
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-cf.log',
- match='"push:eth0" eth0',
- desc='nrlsmf-cf.log contains group "push:eth0" eth0',
-)
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-cf.log',
- match='"push:eth1" eth1',
- desc='nrlsmf-cf.log contains group "push:eth1" eth1',
-)
-
-wait_step(
- "h2",
- "tail -n1 iperf-server.log",
- match="8 pps",
- desc="Receiving 239.0.0.1 at full rate of 8 pps",
-)
-
-# Elastic flooding
-step("r1", "pkill nrlsmf")
-
-wait_step(
- "r1",
- 'pgrep -af "nrlsmf"',
- match="",
- desc="stopped nrlsmf",
-)
-section("Start nrlsmf with elastic flooding r1 ")
-step(
- "r1",
- "nrlsmf debug 4 add net,cf,eth0,eth1 elastic net &> nrlsmf-elastic.log &",
-)
-
-wait_step(
- "r1",
- 'pgrep -af "nrlsmf.*elastic net"',
- match="elastic net",
- desc="nrlsmf is started with elastic group net",
-)
-
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-cf.log',
- match='"net" eth0,eth1',
- desc='nrlsmf-elastic.log contains group "net" eth0,eth1',
-)
-
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-cf.log',
- match='"push:eth0" eth0',
- desc='nrlsmf-elastic.log contains group "push:eth0" eth0',
-)
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-cf.log',
- match='"push:eth1" eth1',
- desc='nrlsmf-elastic.log contains group "push:eth1" eth1',
-)
-# In Elastic flooding, the flow is rate limited to 1.00 KBytes per second.
-wait_step(
- "h2",
- "tail -n4 iperf-server.log",
- match="1.00 KBytes",
- desc="Receiving 239.0.0.1 rate limited to 1 pps",
- timeout=30,
-)
-
-# Advertise mode
-step("r1", "pkill nrlsmf")
-
-wait_step(
- "r1",
- 'pgrep -af "nrlsmf"',
- match="",
- desc="stopped nrlsmf",
-)
-section("Start nrlsmf with advertise mode r1 ")
-step(
- "r1",
- "nrlsmf debug 4 advertise add net,cf,eth0,eth1 elastic net &> nrlsmf-advertise.log &",
-)
-
-wait_step(
- "r1",
- 'pgrep -af "nrlsmf.*advertise.*elastic net"',
- match="advertise",
- desc="nrlsmf is started with advertise + elastic group net",
-)
-
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-advertise.log',
- match='"net" eth0,eth1',
- desc='nrlsmf-advertise.log contains group "net" eth0,eth1',
-)
-
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-advertise.log',
- match='"push:eth0" eth0',
- desc='nrlsmf-advertise.log contains group "push:eth0" eth0',
-)
-wait_step(
- "r1",
- 'grep "regular group" nrlsmf-advertise.log',
- match='"push:eth1" eth1',
- desc='nrlsmf-advertise.log contains group "push:eth1" eth1',
-)
-
-step(
- "h2",
- "tcpdump -lnni eth0 'udp dst port 5555 and dst 224.0.0.55' &> tcpdump-emadv.log &",
-)
-
-wait_step(
- "h2",
- "grep -m1 '224.0.0.55.5555' tcpdump-emadv.log",
- match="224.0.0.55.5555",
- desc="EM_ADV packets seen on h2 eth0 in advertise mode",
- timeout=30,
-)
diff --git a/tests/mutests/1hop_smf/mutest_smf_advertise.py b/tests/mutests/1hop_smf/mutest_smf_advertise.py
new file mode 100644
index 0000000..0f32bf5
--- /dev/null
+++ b/tests/mutests/1hop_smf/mutest_smf_advertise.py
@@ -0,0 +1,154 @@
+"""Example: nrlsmf Elastic Multicast advertise mode (EM_ADV control messages).
+
+Topology (shared with other tests in this directory):
+
+ h0 -- lan0 --\\
+ r0
+ h1 -- lan1 --/
+
+What "advertise" means here
+-------------------------------
+mutest_smf_elastic.py showed Elastic Multicast's data-plane behavior
+(rate-limited delivery instead of blind flooding). This test checks its
+control plane instead: with `advertise` enabled alongside `elastic`,
+nrlsmf periodically emits EM_ADV ("advertisement") messages -- sent to
+the well-known Elastic Multicast control address 224.0.0.55, UDP port
+5555 -- announcing multicast flow/reachability information so
+downstream EM-aware nodes can discover and route toward active flows.
+This test starts the same h0 -> 239.0.0.1 flow so EM has an active
+flow to announce, then captures EM_ADV packets with tcpdump on h1's
+LAN.
+
+What this example covers
+-------------------------
+* nrlsmf started with `advertise add net,cf,eth0,eth1 elastic net` on
+ r0 -- the same EM-enabled group as mutest_smf_elastic.py, with
+ advertisement enabled.
+* An iperf UDP multicast flow from h0 so EM has an active flow to
+ advertise.
+* tcpdump on h1 capturing EM_ADV packets (UDP, destination
+ 224.0.0.55:5555) to confirm the control-plane advertisement traffic
+ is actually being sent.
+
+See mutest_smf_cli.py (CLI sanity checks, no traffic), mutest_smf_merge.py
+(forced two-interface relay), mutest_smf_cf.py (classical flooding),
+and mutest_smf_elastic.py (EM data-plane rate limiting) for the other
+nrlsmf modes on this same topology.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from onehop_hosts import cleanup_iperf
+from onehop_hosts import setup_mcast_route
+from onehop_hosts import start_mcast_client
+from onehop_hosts import start_mcast_server
+from smf_cli import check_common_show
+from smf_cli import check_show_groups
+
+EM_ADV_ADDR = "224.0.0.55"
+EM_ADV_PORT = "5555"
+
+section("Verify interfaces are ready")
+
+for node in ("h0", "h1", "r0"):
+ step(node, "ethtool -K eth0 rx off tx off")
+
+step("r0", "ethtool -K eth1 rx off tx off")
+
+wait_step(
+ "h0",
+ "ip -br addr show dev eth0",
+ match="10.0.0.2/24",
+ desc="h0 has address 10.0.0.2/24",
+)
+wait_step(
+ "h1",
+ "ip -br addr show dev eth0",
+ match="10.0.1.2/24",
+ desc="h1 has address 10.0.1.2/24",
+)
+
+section("Start nrlsmf with Elastic Multicast advertise mode on r0")
+
+step(
+ "r0",
+ "nrlsmf debug 4 advertise add net,cf,eth0,eth1 elastic net "
+ "&> nrlsmf-advertise.log &",
+)
+
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf.*advertise.*elastic net"',
+ match="advertise",
+ desc="nrlsmf is started with advertise + elastic group net",
+)
+
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-advertise.log',
+ match='"net" eth0,eth1',
+ desc='nrlsmf-advertise.log contains group "net" eth0,eth1',
+)
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-advertise.log',
+ match='"push:eth0" eth0',
+ desc='nrlsmf-advertise.log contains the implicit "push:eth0" sub-group',
+)
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-advertise.log',
+ match='"push:eth1" eth1',
+ desc='nrlsmf-advertise.log contains the implicit "push:eth1" sub-group',
+)
+
+section("nrlsmf --cli show commands (json)")
+
+check_common_show("r0", group_name="net", ifaces=("eth0", "eth1"))
+
+section("Capture EM_ADV control messages on h1")
+
+step(
+ "h1",
+ f"tcpdump -lnni eth0 'udp dst port {EM_ADV_PORT} and dst {EM_ADV_ADDR}' "
+ "&> tcpdump-emadv.log &",
+)
+
+setup_mcast_route(step)
+start_mcast_server(step)
+start_mcast_client(step, wait_step)
+
+wait_step(
+ "h1",
+ f"grep -m1 '{EM_ADV_ADDR}.{EM_ADV_PORT}' tcpdump-emadv.log",
+ match=f"{EM_ADV_ADDR}.{EM_ADV_PORT}",
+ desc="EM_ADV packets seen on h1 eth0 in advertise mode",
+ timeout=30,
+)
+
+section("nrlsmf --cli show groups json (advertised EM flow)")
+
+check_show_groups("r0", mcast_addr="239.0.0.1")
+
+section("Cleanup")
+
+cleanup_iperf(step)
+step("h1", "pkill tcpdump || true")
+step("r0", "pkill nrlsmf || true")
+
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf" || true',
+ match="",
+ desc="r0 nrlsmf stopped",
+)
+
+test_step(True, "nrlsmf advertise mode mutest completed")
diff --git a/tests/mutests/1hop_smf/mutest_smf_cf.py b/tests/mutests/1hop_smf/mutest_smf_cf.py
new file mode 100644
index 0000000..959ba1f
--- /dev/null
+++ b/tests/mutests/1hop_smf/mutest_smf_cf.py
@@ -0,0 +1,135 @@
+"""Example: nrlsmf classical flooding (`cf`) between two host LANs.
+
+Topology (shared with other tests in this directory):
+
+ h0 -- lan0 --\\
+ r0
+ h1 -- lan1 --/
+
+What "classical flooding" means here
+---------------------------------------
+`cf ` (Classical Flooding) is SMF's baseline relay
+algorithm: flood received multicast out every other interface in the
+group, with duplicate-packet detection to avoid retransmission storms.
+Here it's invoked via the more general `add ,cf,`
+form, which creates a named interface group ("net") using the `cf`
+relay algorithm.
+
+One detail worth calling out: nrlsmf's log shows not just the "net"
+group itself but also two automatically-created "push:eth0" and
+"push:eth1" sub-groups -- one per member interface. This reflects how
+nrlsmf's internal bookkeeping decomposes a named flooding group into
+per-interface push relationships; it's not something you configure
+directly here; it's what `cf` sets up for you.
+
+What this example covers
+-------------------------
+* nrlsmf started with `add net,cf,eth0,eth1` on r0, and the implicit
+ push:eth0 / push:eth1 sub-groups that come with it.
+* The same multicast reachability check as mutest_smf_merge.py (h0 ->
+ 239.0.0.1 -> h1), confirming `cf` relays traffic just as `merge` did,
+ as expected for this simple two-interface case.
+
+See mutest_smf_cli.py (CLI sanity checks, no traffic), mutest_smf_merge.py
+(forced two-interface relay), mutest_smf_elastic.py (elastic/rate-limited
+multicast), and mutest_smf_advertise.py (EM_ADV control messages) for
+the other nrlsmf modes on this same topology.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from onehop_hosts import cleanup_iperf
+from onehop_hosts import setup_mcast_route
+from onehop_hosts import start_mcast_client
+from onehop_hosts import start_mcast_server
+from onehop_hosts import wait_mcast_receiver
+from smf_cli import check_common_show
+
+MCAST_GROUP = "239.0.0.1"
+
+section("Verify interfaces are ready")
+
+for node in ("h0", "h1", "r0"):
+ step(node, "ethtool -K eth0 rx off tx off")
+
+step("r0", "ethtool -K eth1 rx off tx off")
+
+wait_step(
+ "h0",
+ "ip -br addr show dev eth0",
+ match="10.0.0.2/24",
+ desc="h0 has address 10.0.0.2/24",
+)
+wait_step(
+ "h1",
+ "ip -br addr show dev eth0",
+ match="10.0.1.2/24",
+ desc="h1 has address 10.0.1.2/24",
+)
+
+section("Start nrlsmf classical flooding on r0")
+
+step("r0", "nrlsmf debug 4 add net,cf,eth0,eth1 &> nrlsmf-cf.log &")
+
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf.*net,cf,eth0,eth1"',
+ match="net,cf,eth0,eth1",
+ desc="nrlsmf is started with classical flooding group",
+)
+
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-cf.log',
+ match='"net" eth0,eth1',
+ desc='nrlsmf-cf.log contains group "net" eth0,eth1',
+)
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-cf.log',
+ match='"push:eth0" eth0',
+ desc='nrlsmf-cf.log contains the implicit "push:eth0" sub-group',
+)
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-cf.log',
+ match='"push:eth1" eth1',
+ desc='nrlsmf-cf.log contains the implicit "push:eth1" sub-group',
+)
+
+section("nrlsmf --cli show commands (json)")
+
+check_common_show("r0", group_name="net", ifaces=("eth0", "eth1"))
+
+section("Multicast from h0 reaches h1 via classical flooding on r0")
+
+setup_mcast_route(step)
+start_mcast_server(step)
+start_mcast_client(step, wait_step)
+wait_mcast_receiver(
+ wait_step,
+ match="8 pps",
+ desc=f"h1 receiving {MCAST_GROUP} at full rate of 8 pps",
+)
+
+section("Cleanup")
+
+cleanup_iperf(step)
+step("r0", "pkill nrlsmf || true")
+
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf" || true',
+ match="",
+ desc="r0 nrlsmf stopped",
+)
+
+test_step(True, "nrlsmf classical flooding mutest completed")
diff --git a/tests/mutests/1hop_smf/mutest_smf_cli.py b/tests/mutests/1hop_smf/mutest_smf_cli.py
new file mode 100644
index 0000000..1c295a9
--- /dev/null
+++ b/tests/mutests/1hop_smf/mutest_smf_cli.py
@@ -0,0 +1,127 @@
+"""Example: nrlsmf command-line parsing sanity checks.
+
+Topology (shared with other tests in this directory):
+
+ h0 -- lan0 --\\
+ r0
+ h1 -- lan1 --/
+
+What this example covers
+-------------------------
+No traffic, no forwarding modes -- just a quick sanity check that
+nrlsmf's command-line parser behaves the way the User's Guide says it
+should: `help` prints usage and lists the documented options, `version`
+(and its abbreviation `ver`) prints a version string, and unknown or
+ambiguous commands fail predictably (usage text, non-zero exit) rather
+than crashing or hanging. This is a fast, host-traffic-free check
+that's useful to run before any of the forwarding-mode tests in this
+directory, which all depend on nrlsmf's command line actually working
+as documented.
+
+See mutest_smf_merge.py, mutest_smf_cf.py, mutest_smf_elastic.py, and
+mutest_smf_advertise.py for the actual forwarding-mode tests, all of
+which run on this same h0/r0/h1 topology.
+"""
+
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+section("Verify interfaces are ready")
+
+for node in ("h0", "h1", "r0"):
+ step(node, "ethtool -K eth0 rx off tx off")
+
+step("r0", "ethtool -K eth1 rx off tx off")
+
+wait_step(
+ "h0",
+ "ip -br addr show dev eth0",
+ match="10.0.0.2/24",
+ desc="h0 has address 10.0.0.2/24",
+)
+wait_step(
+ "h1",
+ "ip -br addr show dev eth0",
+ match="10.0.1.2/24",
+ desc="h1 has address 10.0.1.2/24",
+)
+wait_step(
+ "r0",
+ "ip -br addr show dev eth0",
+ match="10.0.0.1/24",
+ desc="r0 eth0 has address 10.0.0.1/24",
+)
+wait_step(
+ "r0",
+ "ip -br addr show dev eth1",
+ match="10.0.1.1/24",
+ desc="r0 eth1 has address 10.0.1.1/24",
+)
+
+section("nrlsmf help")
+
+help_output = step("r0", "sh -lc 'nrlsmf help; echo EXIT:$?'")
+test_step("Usage: nrlsmf" in help_output, "nrlsmf help prints usage", target="r0")
+test_step("EXIT:0" in help_output, "nrlsmf help exits successfully", target="r0")
+test_step(
+ "forward {on | off}" in help_output,
+ "nrlsmf help lists forward option",
+ target="r0",
+)
+test_step(
+ "relay {on | off}" in help_output,
+ "nrlsmf help lists relay option",
+ target="r0",
+)
+test_step(
+ "resequence {on | off}" in help_output,
+ "nrlsmf help lists resequence option",
+ target="r0",
+)
+test_step(
+ "window {on | off}" in help_output,
+ "nrlsmf help lists window option",
+ target="r0",
+)
+
+section("nrlsmf version")
+
+version_output = step("r0", "nrlsmf version")
+test_step(
+ bool(version_output.strip()),
+ "nrlsmf version prints non-empty output",
+ target="r0",
+)
+
+version_abbrev_output = step("r0", "sh -lc 'nrlsmf ver; echo EXIT:$?'")
+test_step(
+ "smf version:" in version_abbrev_output,
+ "abbreviated 'ver' prints version",
+ target="r0",
+)
+test_step("EXIT:0" in version_abbrev_output, "abbreviated 'ver' exits successfully", target="r0")
+
+section("Unknown and ambiguous commands fail predictably")
+
+invalid_cmd_output = step("r0", "sh -lc 'nrlsmf nope; echo EXIT:$?'")
+test_step("Usage: nrlsmf" in invalid_cmd_output, "invalid command prints usage", target="r0")
+test_step("EXIT:0" not in invalid_cmd_output, "invalid command exits non-zero", target="r0")
+
+ambiguous_cmd_output = step("r0", "sh -lc 'nrlsmf r; echo EXIT:$?'")
+test_step("Usage: nrlsmf" in ambiguous_cmd_output, "ambiguous command prints usage", target="r0")
+test_step("EXIT:0" not in ambiguous_cmd_output, "ambiguous command exits non-zero", target="r0")
+
+section("nrlsmf --cli local help (no running daemon)")
+
+cli_help = step("r0", "sh -lc 'nrlsmf --cli -h; echo EXIT:$?'")
+test_step("Usage: nrlsmf --cli" in cli_help, "nrlsmf --cli -h prints usage", target="r0")
+test_step("EXIT:0" in cli_help, "nrlsmf --cli -h exits successfully", target="r0")
+
+cli_show_help = step("r0", "sh -lc 'nrlsmf --cli ?; echo EXIT:$?'")
+test_step("show statistics" in cli_show_help, "nrlsmf --cli ? lists show statistics", target="r0")
+test_step("show tunnel" in cli_show_help, "nrlsmf --cli ? lists show tunnel", target="r0")
+test_step("EXIT:0" in cli_show_help, "nrlsmf --cli ? exits successfully", target="r0")
+
+test_step(True, "nrlsmf CLI sanity checks completed")
diff --git a/tests/mutests/1hop_smf/mutest_smf_elastic.py b/tests/mutests/1hop_smf/mutest_smf_elastic.py
new file mode 100644
index 0000000..e6c333b
--- /dev/null
+++ b/tests/mutests/1hop_smf/mutest_smf_elastic.py
@@ -0,0 +1,148 @@
+"""Example: nrlsmf Elastic Multicast (EM) routing between two host LANs.
+
+Topology (shared with other tests in this directory):
+
+ h0 -- lan0 --\\
+ r0
+ h1 -- lan1 --/
+
+What "elastic" means here
+----------------------------
+`elastic ` overlays Elastic Multicast (EM) routing on top of an
+existing flooding group -- here, the same `cf` group ("net") used in
+mutest_smf_cf.py. Instead of nrlsmf blindly flooding every multicast
+packet it sees, EM manages flows more deliberately: rather than
+matching iperf's full sending rate, this build's EM flow control uses a
+token-bucket that limits an established flow. Where mutest_smf_cf.py's
+classical flooding relayed h0's full 8 pps sending rate through to h1
+unchanged, this test sends the exact same 8 pps flow and expects it to
+arrive at h1 rate-limited to 1 KByte/sec instead -- direct evidence
+that EM's flow control, not blind flooding, is what's forwarding this
+traffic.
+
+What this example covers
+-------------------------
+* nrlsmf started with `add net,cf,eth0,eth1 elastic net` on r0 --
+ the same `cf` group as mutest_smf_cf.py, with EM enabled on it.
+* The same "net" / "push:eth0" / "push:eth1" group log lines as the
+ plain `cf` case (EM sits on top of the same group structure).
+* The same h0 -> 239.0.0.1 -> h1 multicast flow as the other tests
+ here, but this time checked for EM's rate-limited delivery instead
+ of the full sending rate.
+
+See mutest_smf_cli.py (CLI sanity checks, no traffic), mutest_smf_merge.py
+(forced two-interface relay), mutest_smf_cf.py (classical flooding,
+same group without EM), and mutest_smf_advertise.py (EM_ADV control
+messages) for the other nrlsmf modes on this same topology.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from onehop_hosts import cleanup_iperf
+from onehop_hosts import setup_mcast_route
+from onehop_hosts import start_mcast_client
+from onehop_hosts import start_mcast_server
+from onehop_hosts import wait_mcast_receiver
+from smf_cli import check_common_show
+from smf_cli import check_show_groups
+
+MCAST_GROUP = "239.0.0.1"
+
+section("Verify interfaces are ready")
+
+for node in ("h0", "h1", "r0"):
+ step(node, "ethtool -K eth0 rx off tx off")
+
+step("r0", "ethtool -K eth1 rx off tx off")
+
+wait_step(
+ "h0",
+ "ip -br addr show dev eth0",
+ match="10.0.0.2/24",
+ desc="h0 has address 10.0.0.2/24",
+)
+wait_step(
+ "h1",
+ "ip -br addr show dev eth0",
+ match="10.0.1.2/24",
+ desc="h1 has address 10.0.1.2/24",
+)
+
+section("Start nrlsmf with Elastic Multicast on r0")
+
+step(
+ "r0",
+ "nrlsmf debug 4 add net,cf,eth0,eth1 elastic net &> nrlsmf-elastic.log &",
+)
+
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf.*elastic net"',
+ match="elastic net",
+ desc="nrlsmf is started with elastic group net",
+)
+
+# Same group structure as plain `cf` (mutest_smf_cf.py) -- EM overlays
+# on top of it rather than replacing it.
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-elastic.log',
+ match='"net" eth0,eth1',
+ desc='nrlsmf-elastic.log contains group "net" eth0,eth1',
+)
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-elastic.log',
+ match='"push:eth0" eth0',
+ desc='nrlsmf-elastic.log contains the implicit "push:eth0" sub-group',
+)
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-elastic.log',
+ match='"push:eth1" eth1',
+ desc='nrlsmf-elastic.log contains the implicit "push:eth1" sub-group',
+)
+
+section("nrlsmf --cli show commands (json)")
+
+check_common_show("r0", group_name="net", ifaces=("eth0", "eth1"))
+
+section("Multicast from h0 is rate-limited by EM before reaching h1")
+
+setup_mcast_route(step)
+start_mcast_server(step)
+start_mcast_client(step, wait_step)
+# Delivery is rate-limited by EM's flow control to 1.00 KBytes/sec
+# rather than matching h0's full 8 pps sending rate.
+wait_mcast_receiver(
+ wait_step,
+ match="1.00 KBytes",
+ desc="h1 receiving 239.0.0.1 rate-limited by EM to 1 pps",
+ timeout=30,
+)
+
+section("nrlsmf --cli show groups json (active EM flow)")
+
+check_show_groups("r0", mcast_addr=MCAST_GROUP)
+
+section("Cleanup")
+
+cleanup_iperf(step)
+step("r0", "pkill nrlsmf || true")
+
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf" || true',
+ match="",
+ desc="r0 nrlsmf stopped",
+)
+
+test_step(True, "nrlsmf Elastic Multicast mutest completed")
diff --git a/tests/mutests/1hop_smf/mutest_smf_merge.py b/tests/mutests/1hop_smf/mutest_smf_merge.py
new file mode 100644
index 0000000..17daefd
--- /dev/null
+++ b/tests/mutests/1hop_smf/mutest_smf_merge.py
@@ -0,0 +1,135 @@
+"""Example: nrlsmf `merge` -- forced relay between two host LANs.
+
+Topology (shared with other tests in this directory):
+
+ h0 -- lan0 --\\
+ r0
+ h1 -- lan1 --/
+
+What `merge` means here
+--------------------------
+`merge ` is a "Gateway Command": it forces relay of packets
+from any interface in the list to all the others, subject to normal
+duplicate-packet detection and TTL limits, but -- unlike classical
+flooding (`cf`, see mutest_smf_cf.py) -- it never retransmits a packet
+back out the interface it arrived on. For a simple two-interface case
+like this one, that distinction doesn't show up in the traffic pattern
+(there's only one "other" interface to relay to either way), but it
+matters in gateway-style setups with more than two interfaces, where
+`cf` would flood back out every interface including ones that clearly
+don't need it.
+
+This test uses `merge eth0,eth1` on r0 to bridge multicast traffic
+between h0's LAN and h1's LAN: h0 sends multicast, r0 relays it across,
+h1 receives it.
+
+What this example covers
+-------------------------
+* nrlsmf started with `merge eth0,eth1` on r0.
+* An iperf UDP multicast flow from h0 to 239.0.0.1, relayed across to
+ h1's LAN by nrlsmf and received there.
+
+See mutest_smf_cli.py (CLI sanity checks, no traffic), mutest_smf_cf.py
+(classical flooding), mutest_smf_elastic.py (elastic/rate-limited
+multicast), and mutest_smf_advertise.py (EM_ADV control messages) for
+the other nrlsmf modes on this same topology.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from onehop_hosts import cleanup_iperf
+from onehop_hosts import setup_mcast_route
+from onehop_hosts import start_mcast_client
+from onehop_hosts import start_mcast_server
+from onehop_hosts import wait_mcast_receiver
+from smf_cli import check_common_show
+from smf_cli import show_json
+
+MCAST_GROUP = "239.0.0.1"
+
+section("Verify interfaces are ready")
+
+for node in ("h0", "h1", "r0"):
+ step(node, "ethtool -K eth0 rx off tx off")
+
+step("r0", "ethtool -K eth1 rx off tx off")
+
+wait_step(
+ "h0",
+ "ip -br addr show dev eth0",
+ match="10.0.0.2/24",
+ desc="h0 has address 10.0.0.2/24",
+)
+wait_step(
+ "h1",
+ "ip -br addr show dev eth0",
+ match="10.0.1.2/24",
+ desc="h1 has address 10.0.1.2/24",
+)
+
+section("Start nrlsmf merge eth0,eth1 on r0")
+
+step("r0", "nrlsmf debug 4 merge eth0,eth1 &> nrlsmf-merge.log &")
+
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf.*merge eth0,eth1"',
+ match="merge eth0,eth1",
+ desc="nrlsmf is started with merge eth0,eth1",
+)
+
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-merge.log',
+ match='"merge" eth0,eth1',
+ desc="nrlsmf-merge.log contains merge group for eth0,eth1",
+)
+
+section("nrlsmf --cli show commands (json)")
+
+check_common_show("r0", group_name="merge", ifaces=("eth0", "eth1"))
+tunnels = show_json("r0", "show tunnel")
+test_step(isinstance(tunnels, list), "r0 show tunnel json is a list", target="r0")
+neighbors = show_json("r0", "show tunnel neighbors")
+test_step(isinstance(neighbors, list) and len(neighbors) == 0,
+ "r0 show tunnel neighbors json is empty (no GRE)", target="r0")
+wait_step(
+ "r0",
+ 'nrlsmf --cli -c "show version json" -c "show statistics json"',
+ match="jsonVersion",
+ desc="nrlsmf --cli multiple -c commands (json)",
+ timeout=10,
+)
+
+section("Multicast from h0 reaches h1 via merge relay on r0")
+
+setup_mcast_route(step)
+start_mcast_server(step)
+start_mcast_client(step, wait_step)
+wait_mcast_receiver(
+ wait_step,
+ match="8 pps",
+ desc=f"h1 receiving {MCAST_GROUP} at full rate of 8 pps",
+)
+
+section("Cleanup")
+
+cleanup_iperf(step)
+step("r0", "pkill nrlsmf || true")
+
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf" || true',
+ match="",
+ desc="r0 nrlsmf stopped",
+)
+
+test_step(True, "nrlsmf merge mutest completed")
diff --git a/tests/mutests/1hop_smf/onehop_hosts.py b/tests/mutests/1hop_smf/onehop_hosts.py
new file mode 100644
index 0000000..c1e5f9d
--- /dev/null
+++ b/tests/mutests/1hop_smf/onehop_hosts.py
@@ -0,0 +1,54 @@
+"""Host multicast helpers for the 1hop_smf mutests.
+
+h0 sends 239.0.0.1; h1 receives. nrlsmf on r0 is the relay.
+"""
+
+HOST_IFACE = "eth0"
+MCAST_GROUP = "239.0.0.1"
+H0_ADDR = "10.0.0.2"
+IPERF_TTL = "16"
+
+
+def setup_mcast_route(step):
+ for name in ("h0", "h1"):
+ step(name, f"ip route replace {MCAST_GROUP}/32 dev {HOST_IFACE}")
+
+
+def start_mcast_server(step):
+ step(
+ "h1",
+ f"iperf -u -T 4 -i 1 -s -e -B {MCAST_GROUP}%{HOST_IFACE} "
+ "> iperf-server.log 2>&1 &",
+ )
+
+
+def start_mcast_client(step, wait_step):
+ step(
+ "h0",
+ f"iperf -u -T {IPERF_TTL} -t 1000 -i 1 -b 8pps -l 1024 -e "
+ f"-B {H0_ADDR} -c {MCAST_GROUP} &> iperf-client.log &",
+ )
+ wait_step(
+ "h0",
+ 'grep "8 pps" iperf-client.log',
+ match="8 pps",
+ desc=f"h0 sending {MCAST_GROUP} at 8 pps",
+ timeout=30,
+ )
+
+
+def wait_mcast_receiver(wait_step, match="8 pps", desc=None, timeout=20):
+ if desc is None:
+ desc = f"h1 receiving {MCAST_GROUP} at {match}"
+ wait_step(
+ "h1",
+ f'grep "{match}" iperf-server.log',
+ match=match,
+ desc=desc,
+ timeout=timeout,
+ )
+
+
+def cleanup_iperf(step):
+ step("h0", "pkill iperf || true")
+ step("h1", "pkill iperf || true")
diff --git a/tests/mutests/1hop_smf/r1/etc.frr/daemons b/tests/mutests/1hop_smf/r0/etc.frr/daemons
similarity index 100%
rename from tests/mutests/1hop_smf/r1/etc.frr/daemons
rename to tests/mutests/1hop_smf/r0/etc.frr/daemons
diff --git a/tests/mutests/1hop_smf/r1/etc.frr/frr.conf b/tests/mutests/1hop_smf/r0/etc.frr/frr.conf
similarity index 78%
rename from tests/mutests/1hop_smf/r1/etc.frr/frr.conf
rename to tests/mutests/1hop_smf/r0/etc.frr/frr.conf
index 89a5e4a..c829cd2 100644
--- a/tests/mutests/1hop_smf/r1/etc.frr/frr.conf
+++ b/tests/mutests/1hop_smf/r0/etc.frr/frr.conf
@@ -1,8 +1,8 @@
log file /var/log/frr/frr.log
!
interface eth0
- ip address 10.0.1.1/24
+ ip address 10.0.0.1/24
!
interface eth1
- ip address 10.0.2.1/24
+ ip address 10.0.1.1/24
!
diff --git a/tests/mutests/1hop_smf/r1/etc.frr/vtysh.conf b/tests/mutests/1hop_smf/r0/etc.frr/vtysh.conf
similarity index 100%
rename from tests/mutests/1hop_smf/r1/etc.frr/vtysh.conf
rename to tests/mutests/1hop_smf/r0/etc.frr/vtysh.conf
diff --git a/tests/mutests/README.md b/tests/mutests/README.md
index cceef58..0ab315f 100644
--- a/tests/mutests/README.md
+++ b/tests/mutests/README.md
@@ -22,4 +22,16 @@ sudo mutest
```bash
sudo mutest mgre_four_peers
sudo mutest 1hop_smf
+sudo mutest mgre_chained_clouds
```
+
+## Suites
+
+| Directory | Covers |
+|-----------|--------|
+| `1hop_smf/` | Single router, two host LANs: basic nrlsmf CLI and forwarding modes (merge, classical flooding, elastic, advertise) |
+| `mgre_four_peers/` | Four routers across a shared underlay: every GRE/mGRE tunnel mode nrlsmf supports (point-to-point, static NBMA mGRE, NHRP-resolved mGRE, multicast-underlay mGRE, external/metadata GRE), one mode per test |
+| `mgre_chained_clouds/` | All five of those modes chained together end to end across five segments, connected only by nrlsmf relaying (never IP routing) |
+
+See each suite's own `README.md` for its topology and the list of
+individual test files within it.
diff --git a/tests/mutests/kernel_compat.py b/tests/mutests/kernel_compat.py
new file mode 100644
index 0000000..e734480
--- /dev/null
+++ b/tests/mutests/kernel_compat.py
@@ -0,0 +1,36 @@
+"""Skip a mutest when the running kernel is older than a given version."""
+
+import os
+
+from munet.mutest.userapi import test_step
+
+
+def kernel_version():
+ release = os.uname().release.split("-", 1)[0]
+ parts = []
+ for token in release.split(".")[:3]:
+ try:
+ parts.append(int(token))
+ except ValueError:
+ parts.append(0)
+ while len(parts) < 3:
+ parts.append(0)
+ return tuple(parts)
+
+
+def min_kernel_version(min_version):
+ """Return True if the caller should return early.
+
+ min_version is a tuple like (5, 0) or (5, 4, 1).
+ """
+ min_version = tuple(min_version) + (0,) * (3 - len(min_version))
+ have = kernel_version()
+ if have >= min_version:
+ return False
+ need_parts = list(min_version)
+ while len(need_parts) > 2 and need_parts[-1] == 0:
+ need_parts.pop()
+ need = ".".join(str(p) for p in need_parts)
+ have_s = ".".join(str(p) for p in have)
+ test_step(True, f"SKIP: needs Linux {need}+; this kernel is {have_s}")
+ return True
diff --git a/tests/mutests/kinds.yaml b/tests/mutests/kinds.yaml
index 31b92c1..b774e50 100644
--- a/tests/mutests/kinds.yaml
+++ b/tests/mutests/kinds.yaml
@@ -1,4 +1,8 @@
kinds:
+ - name: host
+ cap-add:
+ - NET_ADMIN
+ - NET_RAW
- name: frr
cap-add:
- NET_ADMIN
diff --git a/tests/mutests/mgre_four_peers/p1/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/A/etc.frr/daemons
similarity index 100%
rename from tests/mutests/mgre_four_peers/p1/etc.frr/daemons
rename to tests/mutests/mgre_chained_clouds/A/etc.frr/daemons
diff --git a/tests/mutests/mgre_chained_clouds/A/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/A/etc.frr/frr.conf
new file mode 100644
index 0000000..c73513f
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/A/etc.frr/frr.conf
@@ -0,0 +1,13 @@
+log file /var/log/frr/frr.log
+!
+! A: SMF router in cloud0 (static NBMA mGRE). eth0 is the underlay
+! toward u0; eth1 is the host LAN toward ha (iperf source).
+!
+ip route 0.0.0.0/0 10.0.0.1
+!
+interface eth0
+ ip address 10.0.0.2/24
+!
+interface eth1
+ ip address 192.168.55.1/24
+!
diff --git a/tests/mutests/mgre_four_peers/p1/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/A/etc.frr/vtysh.conf
similarity index 100%
rename from tests/mutests/mgre_four_peers/p1/etc.frr/vtysh.conf
rename to tests/mutests/mgre_chained_clouds/A/etc.frr/vtysh.conf
diff --git a/tests/mutests/mgre_four_peers/p2/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/A2/etc.frr/daemons
similarity index 100%
rename from tests/mutests/mgre_four_peers/p2/etc.frr/daemons
rename to tests/mutests/mgre_chained_clouds/A2/etc.frr/daemons
diff --git a/tests/mutests/mgre_chained_clouds/A2/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/A2/etc.frr/frr.conf
new file mode 100644
index 0000000..afb631c
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/A2/etc.frr/frr.conf
@@ -0,0 +1,13 @@
+log file /var/log/frr/frr.log
+!
+! A2: SMF router in cloud0 (static NBMA mGRE). eth0 is the underlay
+! toward u0; eth1 is the host LAN toward ha2 (iperf receiver).
+!
+ip route 0.0.0.0/0 10.0.1.1
+!
+interface eth0
+ ip address 10.0.1.2/24
+!
+interface eth1
+ ip address 192.168.56.1/24
+!
diff --git a/tests/mutests/mgre_four_peers/p2/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/A2/etc.frr/vtysh.conf
similarity index 100%
rename from tests/mutests/mgre_four_peers/p2/etc.frr/vtysh.conf
rename to tests/mutests/mgre_chained_clouds/A2/etc.frr/vtysh.conf
diff --git a/tests/mutests/mgre_chained_clouds/B/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/B/etc.frr/daemons
new file mode 100644
index 0000000..c14e8db
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/B/etc.frr/daemons
@@ -0,0 +1 @@
+nhrpd=yes
diff --git a/tests/mutests/mgre_chained_clouds/B/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/B/etc.frr/frr.conf
new file mode 100644
index 0000000..bc2ec1d
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/B/etc.frr/frr.conf
@@ -0,0 +1,18 @@
+log file /var/log/frr/frr.log
+!
+! B: SMF router between cloud0 (via u0) and cloud1
+! (via u1). No default route, and no route between cloud0's and
+! cloud1's addressing -- each interface only routes into its own
+! segment's underlay. The only thing connecting cloud0 traffic to
+! cloud1 is nrlsmf running in classical-flooding (`cf`) mode across
+! both of B's overlay (GRE) interfaces -- never IP routing.
+!
+ip route 10.0.0.0/16 10.0.2.1
+ip route 10.1.0.0/16 10.1.0.1
+!
+interface eth0
+ ip address 10.0.2.2/24
+!
+interface eth1
+ ip address 10.1.0.2/24
+!
diff --git a/tests/mutests/mgre_four_peers/p3/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/B/etc.frr/vtysh.conf
similarity index 100%
rename from tests/mutests/mgre_four_peers/p3/etc.frr/vtysh.conf
rename to tests/mutests/mgre_chained_clouds/B/etc.frr/vtysh.conf
diff --git a/tests/mutests/mgre_chained_clouds/B2/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/B2/etc.frr/daemons
new file mode 100644
index 0000000..c14e8db
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/B2/etc.frr/daemons
@@ -0,0 +1 @@
+nhrpd=yes
diff --git a/tests/mutests/mgre_chained_clouds/B2/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/B2/etc.frr/frr.conf
new file mode 100644
index 0000000..f6a5462
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/B2/etc.frr/frr.conf
@@ -0,0 +1,15 @@
+log file /var/log/frr/frr.log
+!
+! B2: SMF router in cloud1 (NHRP-resolved mGRE) and cloud1's hub /
+! NHRP Server (NHS). eth0 is the underlay toward u1; eth1 is the
+! host LAN toward hb2. An operator-controlled peer runs the NHS,
+! never the underlay (u1).
+!
+ip route 0.0.0.0/0 10.1.1.1
+!
+interface eth0
+ ip address 10.1.1.2/24
+!
+interface eth1
+ ip address 192.168.57.1/24
+!
diff --git a/tests/mutests/mgre_four_peers/p4/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/B2/etc.frr/vtysh.conf
similarity index 100%
rename from tests/mutests/mgre_four_peers/p4/etc.frr/vtysh.conf
rename to tests/mutests/mgre_chained_clouds/B2/etc.frr/vtysh.conf
diff --git a/tests/mutests/mgre_chained_clouds/C/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/C/etc.frr/daemons
new file mode 100644
index 0000000..c14e8db
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/C/etc.frr/daemons
@@ -0,0 +1 @@
+nhrpd=yes
diff --git a/tests/mutests/mgre_chained_clouds/C/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/C/etc.frr/frr.conf
new file mode 100644
index 0000000..c3ee096
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/C/etc.frr/frr.conf
@@ -0,0 +1,16 @@
+log file /var/log/frr/frr.log
+!
+! C: SMF router between cloud1 (via u1) and cloud2
+! (via u2). Same principle as B: no route between cloud1's and
+! cloud2's addressing exists here -- only nrlsmf `cf` across C's two
+! overlay interfaces connects the two.
+!
+ip route 10.1.0.0/16 10.1.2.1
+ip route 10.2.0.0/16 10.2.0.1
+!
+interface eth0
+ ip address 10.1.2.2/24
+!
+interface eth1
+ ip address 10.2.0.2/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/C/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/C/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/C/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_four_peers/p3/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/C2/etc.frr/daemons
similarity index 100%
rename from tests/mutests/mgre_four_peers/p3/etc.frr/daemons
rename to tests/mutests/mgre_chained_clouds/C2/etc.frr/daemons
diff --git a/tests/mutests/mgre_chained_clouds/C2/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/C2/etc.frr/frr.conf
new file mode 100644
index 0000000..b9cb2bb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/C2/etc.frr/frr.conf
@@ -0,0 +1,13 @@
+log file /var/log/frr/frr.log
+!
+! C2: leaf peer in cloud2 (multicast-underlay mGRE). eth0 is the
+! underlay toward u2; eth1 is the host LAN toward hc2.
+!
+ip route 0.0.0.0/0 10.2.1.1
+!
+interface eth0
+ ip address 10.2.1.2/24
+!
+interface eth1
+ ip address 192.168.58.1/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/C2/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/C2/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/C2/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_four_peers/p4/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/D/etc.frr/daemons
similarity index 100%
rename from tests/mutests/mgre_four_peers/p4/etc.frr/daemons
rename to tests/mutests/mgre_chained_clouds/D/etc.frr/daemons
diff --git a/tests/mutests/mgre_chained_clouds/D/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/D/etc.frr/frr.conf
new file mode 100644
index 0000000..46cbbb6
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/D/etc.frr/frr.conf
@@ -0,0 +1,16 @@
+log file /var/log/frr/frr.log
+!
+! D: SMF router between cloud2 (via u2) and
+! cloud3 (via u3). Same principle as B and C:
+! no route between cloud2's and cloud3's addressing -- only nrlsmf `cf`
+! across D's two overlay interfaces connects the two.
+!
+ip route 10.2.0.0/16 10.2.2.1
+ip route 10.3.0.0/16 10.3.0.1
+!
+interface eth0
+ ip address 10.2.2.2/24
+!
+interface eth1
+ ip address 10.3.0.2/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/D/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/D/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/D/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_chained_clouds/E/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/E/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_chained_clouds/E/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/E/etc.frr/frr.conf
new file mode 100644
index 0000000..8c30b1c
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/E/etc.frr/frr.conf
@@ -0,0 +1,16 @@
+log file /var/log/frr/frr.log
+!
+! E: SMF router between cloud3 (via u3) and
+! cloud4 (via u4). Same principle throughout this
+! topology: no route between cloud3's and cloud4's addressing -- only
+! nrlsmf `cf` across E's two overlay interfaces connects the two.
+!
+ip route 10.3.0.0/16 10.3.1.1
+ip route 10.4.0.0/16 10.4.0.1
+!
+interface eth0
+ ip address 10.3.1.2/24
+!
+interface eth1
+ ip address 10.4.0.2/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/E/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/E/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/E/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_chained_clouds/E2/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/E2/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_chained_clouds/E2/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/E2/etc.frr/frr.conf
new file mode 100644
index 0000000..0aeb706
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/E2/etc.frr/frr.conf
@@ -0,0 +1,9 @@
+log file /var/log/frr/frr.log
+!
+! E2: leaf peer in cloud4 (external/metadata GRE).
+!
+ip route 0.0.0.0/0 10.4.1.1
+!
+interface eth0
+ ip address 10.4.1.2/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/E2/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/E2/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/E2/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_chained_clouds/F/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/F/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_chained_clouds/F/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/F/etc.frr/frr.conf
new file mode 100644
index 0000000..2a9abc3
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/F/etc.frr/frr.conf
@@ -0,0 +1,14 @@
+log file /var/log/frr/frr.log
+!
+! F: leaf peer in cloud4 (external/metadata GRE). eth0 is the
+! underlay toward u4; eth1 is the host LAN toward hf (far-end
+! iperf receiver).
+!
+ip route 0.0.0.0/0 10.4.2.1
+!
+interface eth0
+ ip address 10.4.2.2/24
+!
+interface eth1
+ ip address 192.168.59.1/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/F/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/F/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/F/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_chained_clouds/README.md b/tests/mutests/mgre_chained_clouds/README.md
new file mode 100644
index 0000000..e27e3cf
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/README.md
@@ -0,0 +1,63 @@
+# Chained-clouds mutest: all five GRE/mGRE modes end to end
+
+One network, five segments in series, each using a different
+GRE/mGRE peer-resolution mode, connected only by nrlsmf relaying
+(never by IP routing):
+
+```
+
+ ha ha2 hb2 hc2 hf
+ | | | | |
+ | A2 B2 C2 E2 |
+ | | | | | |
+ A -- u0 -- B -- u1 -- C -- u2 -- D -- u3 -- E -- u4 -- F
+
+```
+
+| Segment | Underlay | Mode | Peers |
+|---------|----------|------|-------|
+| cloud0 | u0 | Static NBMA mGRE | A, A2, B |
+| cloud1 | u1 | NHRP-resolved mGRE | B, B2 (hub/NHS), C |
+| cloud2 | u2 | Multicast-underlay mGRE | C, C2, D |
+| cloud3 | u3 | Point-to-point GRE | D, E |
+| cloud4 | u4 | External (metadata) GRE | E, E2, F |
+
+Application multicast is sourced on `ha` (off `A`) and received on
+`ha2` (cloud0), `hb2` (cloud1), `hc2` (cloud2), and `hf` (cloud4).
+Each SMF router with a host CFs its host LAN plus its GRE iface so
+nrlsmf is first hop onto the overlay and last hop off it. `B`, `C`,
+`D`, and `E` CF both overlay ifaces. Overlay GRE ifaces are `layered`
+except cloud1's NHRP hub `B2`: nhrpd only programs spoke→hub and
+hub→spokes, so `B2` must replicate overlay multicast onto `gre_c1`.
+`u2` runs `rmerge` only as a stand-in for underlay multicast routing (PIM).
+
+`u0` through `u4` are five separate, disconnected underlay routers —
+there is no IP route anywhere in this topology from one segment's
+addressing to the next. A multicast packet from `ha` can only reach
+`hf` by being relayed hop-by-hop through nrlsmf on `A`, then `B`,
+then `C`, then `D`, then `E`.
+
+Overlay-multicast inject dests are mixed: `map …,dynamic` on A
+(cloud0) and on B/B2/C (cloud1; B2 is the hub replicator). Explicit
+unicast `map`s on A2/B (cloud0) and every cloud4 node. cloud2 uses
+`ujoin`; cloud3 (P2P) needs no `map`.
+
+## Why this exists
+
+Every mode in `../mgre_four_peers/` is tested in isolation. This
+topology chains all five together the way a real deployment might —
+e.g. a MANET island (static NBMA), gatewayed through a DMVPN-style WAN
+(NHRP), onward through a satellite hop (multicast-underlay), and into
+an SDN-managed segment (external/metadata GRE) — and proves one
+multicast flow can cross all of them.
+
+## Run
+
+From `tests/mutests` (requires root, FRR with `nhrpd` enabled, `nrlsmf`
+on PATH):
+
+```bash
+sudo mutest mgre_chained_clouds
+```
+
+Skipped on Linux < 5.0 (collect-md last segment).
diff --git a/tests/mutests/mgre_chained_clouds/chained_hosts.py b/tests/mutests/mgre_chained_clouds/chained_hosts.py
new file mode 100644
index 0000000..3d58aba
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/chained_hosts.py
@@ -0,0 +1,99 @@
+"""Host-LAN helpers for the chained-clouds GRE/mGRE mutest.
+
+Iperf is sourced on ha (off A) and received on ha2/hb2/hc2/hf so
+nrlsmf CF is the first hop onto the overlay and the last hop off it
+on each cloud.
+"""
+
+HOST_IFACE = "eth0"
+ROUTER_HOST_IFACE = "eth1"
+IPERF_TTL = "16"
+
+# (router, router eth1 addr, host name, host addr)
+HOST_LANS = (
+ ("A", "192.168.55.1", "ha", "192.168.55.2"),
+ ("A2", "192.168.56.1", "ha2", "192.168.56.2"),
+ ("B2", "192.168.57.1", "hb2", "192.168.57.2"),
+ ("C2", "192.168.58.1", "hc2", "192.168.58.2"),
+ ("F", "192.168.59.1", "hf", "192.168.59.2"),
+)
+
+SOURCE_HOST = "ha"
+SOURCE_HOST_ADDR = "192.168.55.2"
+RECV_HOSTS = ("ha2", "hb2", "hc2", "hf")
+
+
+def setup_host_lan(step, wait_step):
+ """Bring up each listed router's host LAN and its application host."""
+ for router, router_addr, host, host_addr in HOST_LANS:
+ step(router, f"ethtool -K {ROUTER_HOST_IFACE} rx off tx off || true")
+ wait_step(
+ router,
+ f"ip -br addr show dev {ROUTER_HOST_IFACE}",
+ match=router_addr,
+ desc=f"{router} {ROUTER_HOST_IFACE} address {router_addr}",
+ timeout=30,
+ )
+ step(router, "sysctl -w net.ipv4.conf.all.mc_forwarding=0 || true")
+
+ step(host, f"ethtool -K {HOST_IFACE} rx off tx off || true")
+ step(host, f"ip addr add {host_addr}/24 dev {HOST_IFACE} || true")
+ step(host, f"ip link set {HOST_IFACE} up")
+ step(host, f"ip route replace 0.0.0.0/0 via {router_addr}")
+ wait_step(
+ host,
+ f"ip -br addr show dev {HOST_IFACE}",
+ match=host_addr,
+ desc=f"{host} {HOST_IFACE} address {host_addr}",
+ timeout=30,
+ )
+ wait_step(
+ host,
+ f"ping -c1 -W2 {router_addr}",
+ match="1 received",
+ desc=f"{host} reaches {router} on the host LAN",
+ timeout=20,
+ )
+
+
+def start_overlay_mcast_servers(step, receivers, mcast):
+ for name in receivers:
+ step(name, f"ip route replace {mcast}/32 dev {HOST_IFACE}")
+ step(
+ name,
+ f"iperf -u -T 4 -i 1 -s -e -B {mcast}%{HOST_IFACE} "
+ f"> iperf-{name}-server.log 2>&1 &",
+ )
+
+
+def start_host_mcast_client(step, wait_step, mcast):
+ step(SOURCE_HOST, f"ip route replace {mcast}/32 dev {HOST_IFACE}")
+ step(
+ SOURCE_HOST,
+ f"iperf -u -T {IPERF_TTL} -t 1000 -i 1 -b 8pps -l 1024 -e "
+ f"-B {SOURCE_HOST_ADDR} -c {mcast} &> iperf-{SOURCE_HOST}-client.log &",
+ )
+ wait_step(
+ SOURCE_HOST,
+ f"tail -n1 iperf-{SOURCE_HOST}-client.log",
+ match="8 pps",
+ desc="ha sending application multicast at 8 pps",
+ timeout=30,
+ )
+
+
+def wait_overlay_mcast_receivers(wait_step, receivers):
+ for name in receivers:
+ wait_step(
+ name,
+ f'grep "8 pps" iperf-{name}-server.log',
+ match="8 pps",
+ desc=f"{name} receiving application multicast at 8 pps",
+ timeout=20,
+ )
+
+
+def cleanup_iperf(step, receivers):
+ step(SOURCE_HOST, "pkill iperf || true")
+ for name in receivers:
+ step(name, "pkill iperf || true")
diff --git a/tests/mutests/mgre_chained_clouds/munet.yaml b/tests/mutests/mgre_chained_clouds/munet.yaml
new file mode 100644
index 0000000..ab0a067
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/munet.yaml
@@ -0,0 +1,154 @@
+# Five independent, disconnected underlay segments chained in series
+# by SMF routers, each segment using a different GRE/mGRE
+# peer-resolution mode -- covering all five modes nrlsmf supports
+# in one end-to-end network.
+#
+#
+# ha ha2 hb2 hc2 hf
+# | | | | |
+# | A2 B2 C2 E2 |
+# | | | | | |
+# A -- u0 -- B -- u1 -- C -- u2 -- D -- u3 -- E -- u4 -- F
+#
+#
+# cloud0 (u0, 3 peers): A, A2, B -- static NBMA mGRE
+# cloud1 (u1, 3 peers): B, B2 (hub/NHS), C -- NHRP-resolved mGRE
+# cloud2 (u2, 3 peers): C, C2, D -- multicast-underlay mGRE
+# cloud3 (u3, 2 peers): D, E -- point-to-point GRE
+# cloud4 (u4, 3 peers): E, E2, F -- external (metadata) GRE
+#
+# u0..u4 are five SEPARATE underlay routers. There is no IP route
+# anywhere in this topology connecting one segment's addressing to
+# the next. B, C, D, and E route only within each attached
+# segment. The only thing connecting cloud0's
+# traffic all the way to cloud4 is nrlsmf CF -- never IP routing.
+#
+# Application multicast is sourced on ha (off A) and received on
+# ha2 / hb2 / hc2 / hf so nrlsmf CF is first hop onto the overlay
+# and last hop off it on each cloud.
+topology:
+ networks:
+ # cloud0 (behind u0) -- static NBMA mGRE
+ - name: lanA
+ - name: lanA2
+ - name: lanB0 # B's cloud0-facing leg
+ # cloud1 (behind u1) -- NHRP-resolved mGRE
+ - name: lanB1 # B's cloud1-facing leg
+ - name: lanB2 # B2 (cloud1 hub/NHS)
+ - name: lanC0 # C's cloud1-facing leg
+ # cloud2 (behind u2) -- multicast-underlay mGRE
+ - name: lanC1 # C's cloud2-facing leg
+ - name: lanC2 # C2
+ - name: lanD0 # D's cloud2-facing leg
+ # cloud3 (behind u3) -- point-to-point GRE
+ - name: lanD1 # D's cloud3-facing leg
+ - name: lanE0 # E's cloud3-facing leg
+ # cloud4 (behind u4) -- external (metadata) GRE
+ - name: lanE1 # E's cloud4-facing leg
+ - name: lanE2 # E2
+ - name: lanF0 # F
+ # host LANs -- iperf source/receivers, not part of any GRE overlay
+ - name: lanha # off A (source)
+ - name: lanha2 # off A2 (cloud0)
+ - name: lanhb2 # off B2 (cloud1)
+ - name: lanhc2 # off C2 (cloud2)
+ - name: lanhf # off F (cloud4)
+ nodes:
+ - name: u0
+ kind: frr
+ connections:
+ - to: lanA
+ - to: lanA2
+ - to: lanB0
+ - name: u1
+ kind: frr
+ connections:
+ - to: lanB1
+ - to: lanB2
+ - to: lanC0
+ - name: u2
+ kind: frr
+ connections:
+ - to: lanC1
+ - to: lanC2
+ - to: lanD0
+ - name: u3
+ kind: frr
+ connections:
+ - to: lanD1
+ - to: lanE0
+ - name: u4
+ kind: frr
+ connections:
+ - to: lanE1
+ - to: lanE2
+ - to: lanF0
+ - name: ha
+ kind: host
+ connections:
+ - to: lanha
+ - name: ha2
+ kind: host
+ connections:
+ - to: lanha2
+ - name: hb2
+ kind: host
+ connections:
+ - to: lanhb2
+ - name: hc2
+ kind: host
+ connections:
+ - to: lanhc2
+ - name: hf
+ kind: host
+ connections:
+ - to: lanhf
+ - name: A
+ kind: frr
+ connections:
+ - to: lanA
+ - to: lanha
+ - name: A2
+ kind: frr
+ connections:
+ - to: lanA2
+ - to: lanha2
+ - name: B
+ kind: frr
+ connections:
+ - to: lanB0 # cloud0 leg
+ - to: lanB1 # cloud1 leg
+ - name: B2
+ kind: frr
+ connections:
+ - to: lanB2
+ - to: lanhb2
+ - name: C
+ kind: frr
+ connections:
+ - to: lanC0 # cloud1 leg
+ - to: lanC1 # cloud2 leg
+ - name: C2
+ kind: frr
+ connections:
+ - to: lanC2
+ - to: lanhc2
+ - name: D
+ kind: frr
+ connections:
+ - to: lanD0 # cloud2 leg
+ - to: lanD1 # cloud3 leg
+ - name: E
+ kind: frr
+ connections:
+ - to: lanE0 # cloud3 leg
+ - to: lanE1 # cloud4 leg
+ - name: E2
+ kind: frr
+ connections:
+ - to: lanE2
+ - name: F
+ kind: frr
+ connections:
+ - to: lanF0
+ - to: lanhf
diff --git a/tests/mutests/mgre_chained_clouds/mutest_chained_clouds.py b/tests/mutests/mgre_chained_clouds/mutest_chained_clouds.py
new file mode 100644
index 0000000..8eb7396
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/mutest_chained_clouds.py
@@ -0,0 +1,569 @@
+"""Example: one chained network exercising all five GRE/mGRE modes.
+
+Topology:
+
+ ha ha2 hb2 hc2 hf
+ | | | | |
+ | A2 B2 C2 E2 |
+ | | | | | |
+ A -- u0 -- B -- u1 -- C -- u2 -- D -- u3 -- E -- u4 -- F
+
+ cloud0 (u0, 3 peers): A, A2, B -- static NBMA mGRE
+ cloud1 (u1, 3 peers): B, B2 (hub/NHS), C -- NHRP-resolved mGRE
+ cloud2 (u2, 3 peers): C, C2, D -- multicast-underlay mGRE
+ cloud3 (u3, 2 peers): D, E -- point-to-point GRE
+ cloud4 (u4, 3 peers): E, E2, F -- external (metadata) GRE
+
+ ha is the iperf source (off A). ha2 / hb2 / hc2 / hf are receivers
+ off A2 / B2 / C2 / F -- one host on each cloud, last hop is nrlsmf
+ CF onto that host LAN.
+
+Why this test exists
+---------------------
+Every other test in tests/mutests/mgre_four_peers/ demonstrates one
+GRE/mGRE mode in isolation. Real deployments chain several of these
+together -- e.g. a MANET island using static NBMA mGRE, gatewayed
+through a DMVPN-style WAN using NHRP, onward through a satellite hop
+using multicast-underlay mGRE, and finally into an SDN-managed segment
+using external/metadata GRE. This test builds exactly that kind of
+chain, five segments end to end, and proves a single multicast flow
+can cross all five.
+
+u0 through u4 are five *separate*, disconnected underlay routers --
+there is no IP route anywhere in this topology from one segment's
+addressing to the next. B, C, D, and E each sit on two adjacent
+clouds, but their frr.conf routes only reach into those clouds. The
+only thing connecting cloud0's traffic all the way to cloud4 is nrlsmf
+running in classical-flooding (`cf`) mode across both of each SMF
+router's overlay interfaces -- never IP routing. u2 uses `rmerge` only
+as a stand-in for underlay multicast routing (PIM), not as an overlay
+SMF router.
+
+What this example covers
+-------------------------
+* Building all five GRE/mGRE modes in a single network.
+* Overlay-multicast inject dests mixed across nodes:
+ - cloud0: A `map …,dynamic`; A2 and B explicit unicast `map`s
+ - cloud1: B and C `map …,dynamic` (spoke→hub); B2 `map …,dynamic`
+ and is not layered (hub replicates to the other spoke)
+ - cloud2: `ujoin` on the underlay iface (single multicast remote)
+ - cloud3: no `map` (one configured remote)
+ - cloud4: explicit unicast `map`s (`dynamic` does not apply)
+* Application multicast sourced on ha, received on ha2/hb2/hc2/hf.
+
+See tests/mutests/mgre_four_peers/ for each of these five modes
+documented and tested individually.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from chained_hosts import RECV_HOSTS
+from chained_hosts import cleanup_iperf
+from chained_hosts import setup_host_lan
+from chained_hosts import start_host_mcast_client
+from chained_hosts import start_overlay_mcast_servers
+from chained_hosts import wait_overlay_mcast_receivers
+from kernel_compat import min_kernel_version
+from smf_cli import check_common_show
+from smf_cli import check_show_neighbors
+from smf_cli import check_show_tunnel
+
+# ---------------------------------------------------------------------
+# cloud0 (u0): static NBMA mGRE -- A, A2, B
+# ---------------------------------------------------------------------
+GRE_C0 = "gre_c0"
+C0_PEERS = {
+ "A": {"underlay": "10.0.0.2", "overlay": "172.16.0.1"},
+ "A2": {"underlay": "10.0.1.2", "overlay": "172.16.0.2"},
+ "B": {"underlay": "10.0.2.2", "overlay": "172.16.0.3"},
+}
+
+# ---------------------------------------------------------------------
+# cloud1 (u1): NHRP-resolved mGRE -- B, B2 (hub/NHS), C
+# ---------------------------------------------------------------------
+GRE_C1 = "gre_c1"
+GRE_KEY_C1 = "42"
+C1_HUB = "B2"
+C1_PEERS = {
+ "B": {"underlay": "10.1.0.2", "overlay": "172.17.0.2"},
+ "B2": {"underlay": "10.1.1.2", "overlay": "172.17.0.1"},
+ "C": {"underlay": "10.1.2.2", "overlay": "172.17.0.3"},
+}
+
+# ---------------------------------------------------------------------
+# cloud2 (u2): multicast-underlay mGRE -- C, C2, D
+# u2 itself runs `rmerge` across its three LANs as a PIM stand-in.
+# ---------------------------------------------------------------------
+GRE_C2 = "gre_c2"
+UNDERLAY_MCAST_C2 = "239.1.1.1"
+C2_PEERS = {
+ "C": {"underlay": "10.2.0.2", "overlay": "172.18.0.1", "uj_iface": "eth1"},
+ "C2": {"underlay": "10.2.1.2", "overlay": "172.18.0.2", "uj_iface": "eth0"},
+ "D": {"underlay": "10.2.2.2", "overlay": "172.18.0.3", "uj_iface": "eth0"},
+}
+U2_IFACES = "eth0,eth1,eth2"
+
+# ---------------------------------------------------------------------
+# cloud3 (u3): point-to-point GRE -- D <-> E
+# ---------------------------------------------------------------------
+GRE_P2P = "gre_p2p"
+P2P_PEERS = {
+ "D": {"underlay": "10.3.0.2", "overlay": "172.19.0.1"},
+ "E": {"underlay": "10.3.1.2", "overlay": "172.19.0.2"},
+}
+
+# ---------------------------------------------------------------------
+# cloud4 (u4): external (metadata) GRE -- E, E2, F
+# ---------------------------------------------------------------------
+GRE_C4 = "gre_c4"
+GRE_KEY_C4 = "300"
+C4_PEERS = {
+ "E": {"underlay": "10.4.0.2", "overlay": "172.20.0.1"},
+ "E2": {"underlay": "10.4.1.2", "overlay": "172.20.0.2"},
+ "F": {"underlay": "10.4.2.2", "overlay": "172.20.0.3"},
+}
+
+OVERLAY_MCAST = "239.0.0.1"
+
+UNDERLAY_ROUTERS = {
+ "u0": {"eth0": "10.0.0.1", "eth1": "10.0.1.1", "eth2": "10.0.2.1"},
+ "u1": {"eth0": "10.1.0.1", "eth1": "10.1.1.1", "eth2": "10.1.2.1"},
+ "u2": {"eth0": "10.2.0.1", "eth1": "10.2.1.1", "eth2": "10.2.2.1"},
+ "u3": {"eth0": "10.3.0.1", "eth1": "10.3.1.1"},
+ "u4": {"eth0": "10.4.0.1", "eth1": "10.4.1.1", "eth2": "10.4.2.1"},
+}
+LEAF_NODES = {
+ "C2": {"eth0": "10.2.1.2"},
+ "E2": {"eth0": "10.4.1.2"},
+ "F": {"eth0": "10.4.2.2"},
+}
+SMF_ROUTERS = {
+ "A": {"eth0": "10.0.0.2"},
+ "A2": {"eth0": "10.0.1.2"},
+ "B": {"eth0": "10.0.2.2", "eth1": "10.1.0.2"},
+ "B2": {"eth0": "10.1.1.2"},
+ "C": {"eth0": "10.1.2.2", "eth1": "10.2.0.2"},
+ "D": {"eth0": "10.2.2.2", "eth1": "10.3.0.2"},
+ "E": {"eth0": "10.3.1.2", "eth1": "10.4.0.2"},
+}
+
+ALL_NODE_ADDRS = {**UNDERLAY_ROUTERS, **LEAF_NODES, **SMF_ROUTERS}
+
+if min_kernel_version((5, 0)):
+ return "skip"
+
+
+section("Disable offloads and wait for underlay addresses")
+
+for name, ifaces in ALL_NODE_ADDRS.items():
+ for ifname, addr in ifaces.items():
+ step(name, f"ethtool -K {ifname} rx off tx off || true")
+ wait_step(
+ name,
+ f"ip -br addr show dev {ifname}",
+ match=addr,
+ desc=f"{name} {ifname} address {addr}",
+ timeout=30,
+ )
+
+for name in UNDERLAY_ROUTERS:
+ step(name, "sysctl -w net.ipv4.ip_forward=1")
+
+setup_host_lan(step, wait_step)
+
+section("Underlay reachability within each segment (never across segments)")
+
+wait_step("A", "ping -c1 -W2 10.0.2.2", match="1 received", desc="A reaches B within cloud0", timeout=20)
+wait_step("B2", "ping -c1 -W2 10.1.0.2", match="1 received", desc="B2 reaches B within cloud1", timeout=20)
+wait_step("C2", "ping -c1 -W2 10.2.2.2", match="1 received", desc="C2 reaches D within cloud2", timeout=20)
+wait_step("D", "ping -c1 -W2 10.3.1.2", match="1 received", desc="D reaches E within cloud3", timeout=20)
+wait_step("E2", "ping -c1 -W2 10.4.2.2", match="1 received", desc="E2 reaches F within cloud4", timeout=20)
+
+
+section("Build cloud0: static NBMA mGRE (A, A2, B)")
+
+for name, cfg in C0_PEERS.items():
+ step(name, "ip addr flush dev gre0 2>/dev/null || true")
+ step(name, "ip link set gre0 down 2>/dev/null || true")
+ step(name, f"ip link del {GRE_C0} 2>/dev/null || true")
+ step(
+ name,
+ f"ip link add name {GRE_C0} type gre "
+ f"local {cfg['underlay']} remote 0.0.0.0 ttl 64",
+ )
+ step(name, f"ip addr add {cfg['overlay']}/24 dev {GRE_C0}")
+ step(name, f"ip link set {GRE_C0} multicast on")
+ step(name, f"ip link set {GRE_C0} up")
+ for other, ocfg in C0_PEERS.items():
+ if other == name:
+ continue
+ step(
+ name,
+ f"ip neigh replace {ocfg['overlay']} lladdr {ocfg['underlay']} "
+ f"nud permanent dev {GRE_C0}",
+ )
+ wait_step(name, f"ip -br link show {GRE_C0}", match="UP", desc=f"{name} {GRE_C0} is UP")
+
+section("Build cloud1: NHRP-resolved mGRE (B, B2=hub/NHS, C)")
+
+for name, cfg in C1_PEERS.items():
+ step(name, "ip addr flush dev gre0 2>/dev/null || true")
+ step(name, "ip link set gre0 down 2>/dev/null || true")
+ step(name, f"ip link del {GRE_C1} 2>/dev/null || true")
+ step(
+ name,
+ f"ip tunnel add {GRE_C1} mode gre "
+ f"local {cfg['underlay']} key {GRE_KEY_C1} ttl 64",
+ )
+ step(name, f"ip addr add {cfg['overlay']}/32 dev {GRE_C1}")
+ step(name, f"ip link set {GRE_C1} multicast on")
+ step(name, f"ip link set {GRE_C1} up")
+ wait_step(name, f"ip -br link show {GRE_C1}", match="UP", desc=f"{name} {GRE_C1} is UP")
+
+wait_step(C1_HUB, "pgrep -af nhrpd", match="nhrpd", desc=f"{C1_HUB} nhrpd is running", timeout=20)
+step(
+ C1_HUB,
+ f"vtysh -c 'configure terminal' "
+ f"-c 'interface {GRE_C1}' "
+ f"-c 'ip nhrp network-id 1' "
+ f"-c 'ip nhrp registration no-unique'",
+)
+
+for name in ("B", "C"):
+ wait_step(name, "pgrep -af nhrpd", match="nhrpd", desc=f"{name} nhrpd is running", timeout=20)
+ step(
+ name,
+ f"vtysh -c 'configure terminal' "
+ f"-c 'interface {GRE_C1}' "
+ f"-c 'ip nhrp network-id 1' "
+ f"-c 'ip nhrp nhs {C1_PEERS[C1_HUB]['overlay']} "
+ f"nbma {C1_PEERS[C1_HUB]['underlay']}' "
+ f"-c 'ip nhrp registration no-unique'",
+ )
+
+for name in ("B", "C"):
+ wait_step(
+ C1_HUB,
+ "vtysh -c 'show ip nhrp cache'",
+ match=C1_PEERS[name]["overlay"],
+ desc=f"{C1_HUB} NHRP cache shows {name} registered",
+ timeout=60,
+ )
+
+section("Build cloud2: multicast-underlay mGRE (C, C2, D) + u2 as PIM stand-in")
+
+step("u2", f"nrlsmf debug 4 instance smf-u2-underlay rmerge {U2_IFACES} &> nrlsmf-u2-underlay.log &")
+wait_step(
+ "u2",
+ 'pgrep -af "nrlsmf.*instance smf-u2-underlay"',
+ match="smf-u2-underlay",
+ desc="u2 underlay nrlsmf (PIM stand-in) running",
+ timeout=20,
+)
+wait_step(
+ "u2",
+ 'grep "regular group" nrlsmf-u2-underlay.log',
+ match="merge",
+ desc="u2 nrlsmf log shows merge group",
+ timeout=20,
+)
+
+for name, cfg in C2_PEERS.items():
+ step(name, "ip addr flush dev gre0 2>/dev/null || true")
+ step(name, "ip link set gre0 down 2>/dev/null || true")
+ step(name, f"ip link del {GRE_C2} 2>/dev/null || true")
+ step(
+ name,
+ f"ip tunnel add {GRE_C2} mode gre "
+ f"local {cfg['underlay']} remote {UNDERLAY_MCAST_C2} ttl 64",
+ )
+ step(name, f"ip addr add {cfg['overlay']}/24 dev {GRE_C2}")
+ step(name, f"ip link set {GRE_C2} multicast on")
+ step(name, f"ip link set {GRE_C2} up")
+ wait_step(name, f"ip -br link show {GRE_C2}", match="UP", desc=f"{name} {GRE_C2} is UP")
+
+section("Build cloud3: point-to-point GRE (D, E)")
+
+for local, remote in (("D", "E"), ("E", "D")):
+ step(local, "ip addr flush dev gre0 2>/dev/null || true")
+ step(local, "ip link set gre0 down 2>/dev/null || true")
+ step(local, f"ip link del {GRE_P2P} 2>/dev/null || true")
+ step(
+ local,
+ f"ip link add name {GRE_P2P} type gre "
+ f"local {P2P_PEERS[local]['underlay']} "
+ f"remote {P2P_PEERS[remote]['underlay']} ttl 64",
+ )
+ step(local, f"ip addr add {P2P_PEERS[local]['overlay']}/30 dev {GRE_P2P}")
+ step(local, f"ip link set {GRE_P2P} multicast on")
+ step(local, f"ip link set {GRE_P2P} up")
+ wait_step(local, f"ip -br link show {GRE_P2P}", match="UP", desc=f"{local} {GRE_P2P} is UP")
+
+section("Build cloud4: external (metadata) GRE (E, E2, F)")
+
+for name, cfg in C4_PEERS.items():
+ step(name, "ip addr flush dev gre0 2>/dev/null || true")
+ step(name, "ip link set gre0 down 2>/dev/null || true")
+ step(name, f"ip link del {GRE_C4} 2>/dev/null || true")
+ step(name, f"ip link add name {GRE_C4} type gre external")
+ step(name, f"ip link set {GRE_C4} multicast on")
+ step(name, f"ip link set {GRE_C4} up")
+ step(name, f"ip addr add {cfg['overlay']}/24 dev {GRE_C4}")
+ for other, ocfg in C4_PEERS.items():
+ if other == name:
+ continue
+ step(
+ name,
+ f"ip route replace {ocfg['overlay']}/32 encap ip id {GRE_KEY_C4} "
+ f"src {cfg['underlay']} dst {ocfg['underlay']} ttl 64 dev {GRE_C4}",
+ )
+ wait_step(name, f"ip -br link show {GRE_C4}", match="UP", desc=f"{name} {GRE_C4} is UP")
+
+section("Overlay unicast reachability within each segment (before nrlsmf)")
+
+wait_step("A", f"ping -c1 -W3 -I {C0_PEERS['A']['overlay']} {C0_PEERS['B']['overlay']}",
+ match="1 received", desc="A reaches B over cloud0 overlay", timeout=20)
+wait_step("B2", f"ping -c1 -W3 -I {C1_PEERS['B2']['overlay']} {C1_PEERS['C']['overlay']}",
+ match="1 received", desc="B2 reaches C over cloud1 overlay", timeout=20)
+wait_step("C2", f"ping -c1 -W3 -I {GRE_C2} {C2_PEERS['D']['overlay']}",
+ match="1 received", desc="C2 reaches D over cloud2 overlay", timeout=20)
+wait_step("D", f"ping -c1 -W3 {P2P_PEERS['E']['overlay']}",
+ match="1 received", desc="D reaches E over cloud3 overlay", timeout=20)
+wait_step("E2", f"ping -c1 -W3 -I {C4_PEERS['E2']['overlay']} {C4_PEERS['F']['overlay']}",
+ match="1 received", desc="E2 reaches F over cloud4 overlay", timeout=20)
+
+section("Start nrlsmf on overlay routers")
+
+# cloud0: A learns inject dests from neigh; A2/B pin them.
+# cloud1: nhrpd programs neigh; B/C (spokes) map dynamic + layered;
+# B2 (hub) map dynamic, not layered (overlay replicator).
+# cloud2: ujoin on the underlay iface.
+# cloud3: one configured remote, no map.
+# cloud4: explicit unicast maps (no neigh to learn).
+step(
+ "A",
+ "nrlsmf debug 4 "
+ "instance smf-A "
+ f"add overlay,cf,eth1,{GRE_C0} "
+ f"layered {GRE_C0} "
+ f"map {GRE_C0},{C0_PEERS['A']['underlay']},dynamic "
+ "&> nrlsmf-A.log &",
+)
+step(
+ "A2",
+ "nrlsmf debug 4 "
+ "instance smf-A2 "
+ f"add overlay,cf,eth1,{GRE_C0} "
+ f"layered {GRE_C0} "
+ f"map {GRE_C0},{C0_PEERS['A2']['underlay']},{C0_PEERS['A']['underlay']} "
+ f"map {GRE_C0},{C0_PEERS['A2']['underlay']},{C0_PEERS['B']['underlay']} "
+ "&> nrlsmf-A2.log &",
+)
+step(
+ "B",
+ "nrlsmf debug 4 "
+ "instance smf-B "
+ f"add overlay,cf,{GRE_C0},{GRE_C1} "
+ f"layered {GRE_C0},{GRE_C1} "
+ f"map {GRE_C0},{C0_PEERS['B']['underlay']},{C0_PEERS['A']['underlay']} "
+ f"map {GRE_C0},{C0_PEERS['B']['underlay']},{C0_PEERS['A2']['underlay']} "
+ f"map {GRE_C1},{C1_PEERS['B']['underlay']},dynamic "
+ "&> nrlsmf-B.log &",
+)
+step(
+ "B2",
+ "nrlsmf debug 4 "
+ "instance smf-B2 "
+ f"add overlay,cf,eth1,{GRE_C1} "
+ f"map {GRE_C1},{C1_PEERS['B2']['underlay']},dynamic "
+ "&> nrlsmf-B2.log &",
+)
+step(
+ "C",
+ "nrlsmf debug 4 "
+ "instance smf-C "
+ f"add overlay,cf,{GRE_C1},{GRE_C2} "
+ f"layered {GRE_C1},{GRE_C2} "
+ f"map {GRE_C1},{C1_PEERS['C']['underlay']},dynamic "
+ f"ujoin {UNDERLAY_MCAST_C2},{C2_PEERS['C']['uj_iface']} "
+ "&> nrlsmf-C.log &",
+)
+step(
+ "C2",
+ "nrlsmf debug 4 "
+ "instance smf-C2 "
+ f"add overlay,cf,eth1,{GRE_C2} "
+ f"layered {GRE_C2} "
+ f"ujoin {UNDERLAY_MCAST_C2},{C2_PEERS['C2']['uj_iface']} "
+ "&> nrlsmf-C2.log &",
+)
+step(
+ "D",
+ "nrlsmf debug 4 "
+ "instance smf-D "
+ f"add overlay,cf,{GRE_C2},{GRE_P2P} "
+ f"layered {GRE_C2},{GRE_P2P} "
+ f"ujoin {UNDERLAY_MCAST_C2},{C2_PEERS['D']['uj_iface']} "
+ "&> nrlsmf-D.log &",
+)
+step(
+ "E",
+ "nrlsmf debug 4 "
+ "instance smf-E "
+ f"add overlay,cf,{GRE_P2P},{GRE_C4} "
+ f"layered {GRE_P2P},{GRE_C4} "
+ f"map {GRE_C4},{C4_PEERS['E']['underlay']},{C4_PEERS['E2']['underlay']} "
+ f"map {GRE_C4},{C4_PEERS['E']['underlay']},{C4_PEERS['F']['underlay']} "
+ "&> nrlsmf-E.log &",
+)
+step(
+ "E2",
+ "nrlsmf debug 4 "
+ "instance smf-E2 "
+ f"add overlay,cf,{GRE_C4} "
+ f"layered {GRE_C4} "
+ f"map {GRE_C4},{C4_PEERS['E2']['underlay']},{C4_PEERS['E']['underlay']} "
+ f"map {GRE_C4},{C4_PEERS['E2']['underlay']},{C4_PEERS['F']['underlay']} "
+ "&> nrlsmf-E2.log &",
+)
+step(
+ "F",
+ "nrlsmf debug 4 "
+ "instance smf-F "
+ f"add overlay,cf,eth1,{GRE_C4} "
+ f"layered {GRE_C4} "
+ f"map {GRE_C4},{C4_PEERS['F']['underlay']},{C4_PEERS['E']['underlay']} "
+ f"map {GRE_C4},{C4_PEERS['F']['underlay']},{C4_PEERS['E2']['underlay']} "
+ "&> nrlsmf-F.log &",
+)
+
+for name in ("A", "A2", "B", "B2", "C", "C2", "D", "E", "E2", "F"):
+ wait_step(
+ name,
+ f'pgrep -af "nrlsmf.*instance smf-{name}"',
+ match=f"smf-{name}",
+ desc=f"{name} nrlsmf running",
+ timeout=20,
+ )
+ wait_step(
+ name,
+ f'grep "regular group" nrlsmf-{name}.log',
+ match="overlay",
+ desc=f"{name} nrlsmf log shows overlay group",
+ timeout=20,
+ )
+
+section("nrlsmf --cli show tunnel / neighbors (json, all five GRE modes)")
+
+# cloud0 static NBMA: A learns from neigh; A2 has explicit maps.
+check_common_show("A", "smf-A", group_name="overlay", ifaces=("eth1", GRE_C0))
+check_show_tunnel(
+ "A", "smf-A", GRE_C0,
+ local=C0_PEERS["A"]["underlay"],
+ remotes=[C0_PEERS["A2"]["underlay"], C0_PEERS["B"]["underlay"]],
+ overlay_ip=C0_PEERS["A"]["overlay"],
+)
+check_show_neighbors(
+ "A", "smf-A", GRE_C0,
+ remotes=[C0_PEERS["A2"]["underlay"], C0_PEERS["B"]["underlay"]],
+ neighbor_ips=[C0_PEERS["A2"]["overlay"], C0_PEERS["B"]["overlay"]],
+ min_count=2,
+)
+check_show_tunnel(
+ "A2", "smf-A2", GRE_C0,
+ local=C0_PEERS["A2"]["underlay"],
+ remotes=[C0_PEERS["A"]["underlay"], C0_PEERS["B"]["underlay"]],
+ overlay_ip=C0_PEERS["A2"]["overlay"],
+ want_c=True,
+)
+check_show_neighbors(
+ "A2", "smf-A2", GRE_C0,
+ remotes=[C0_PEERS["A"]["underlay"], C0_PEERS["B"]["underlay"]],
+ neighbor_ips=[C0_PEERS["A"]["overlay"], C0_PEERS["B"]["overlay"]],
+ min_count=2,
+ want_c=True,
+)
+
+# cloud1 NHRP: hub B2 sees both spokes.
+check_show_neighbors(
+ "B2", "smf-B2", GRE_C1,
+ remotes=[C1_PEERS["B"]["underlay"], C1_PEERS["C"]["underlay"]],
+ neighbor_ips=[C1_PEERS["B"]["overlay"], C1_PEERS["C"]["overlay"]],
+ min_count=2,
+)
+check_show_neighbors(
+ "B", "smf-B", GRE_C1,
+ remotes=[C1_PEERS["B2"]["underlay"]],
+ neighbor_ips=[C1_PEERS["B2"]["overlay"]],
+ min_count=1,
+)
+
+# cloud2 multicast-underlay remote.
+check_show_tunnel(
+ "C2", "smf-C2", GRE_C2,
+ local=C2_PEERS["C2"]["underlay"],
+ remotes=[UNDERLAY_MCAST_C2],
+ overlay_ip=C2_PEERS["C2"]["overlay"],
+ want_c=False,
+)
+
+# cloud3 P2P: kernel-learned single remote, no map.
+check_show_tunnel(
+ "D", "smf-D", GRE_P2P,
+ local=P2P_PEERS["D"]["underlay"],
+ remotes=[P2P_PEERS["E"]["underlay"]],
+ overlay_ip=P2P_PEERS["D"]["overlay"],
+ want_c=False,
+)
+check_show_neighbors(
+ "D", "smf-D", GRE_P2P,
+ remotes=[P2P_PEERS["E"]["underlay"]],
+ min_count=1,
+)
+
+# cloud4 external GRE: explicit maps, no neigh table.
+check_show_tunnel(
+ "F", "smf-F", GRE_C4,
+ local=C4_PEERS["F"]["underlay"],
+ remotes=[C4_PEERS["E"]["underlay"], C4_PEERS["E2"]["underlay"]],
+ overlay_ip=C4_PEERS["F"]["overlay"],
+ want_c=True,
+)
+check_show_neighbors(
+ "F", "smf-F", GRE_C4,
+ remotes=[C4_PEERS["E"]["underlay"], C4_PEERS["E2"]["underlay"]],
+ min_count=2,
+ want_c=True,
+)
+
+section("Overlay multicast: ha -> ha2 / hb2 / hc2 / hf")
+
+start_overlay_mcast_servers(step, RECV_HOSTS, OVERLAY_MCAST)
+start_host_mcast_client(step, wait_step, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, RECV_HOSTS)
+
+section("Cleanup")
+
+cleanup_iperf(step, RECV_HOSTS)
+
+for name in (
+ "A", "A2", "B", "B2", "C", "C2", "D", "E", "E2", "F", "u2",
+):
+ step(name, "pkill nrlsmf || true")
+ wait_step(
+ name,
+ 'pgrep -af "nrlsmf" || true',
+ match="",
+ desc=f"{name} nrlsmf stopped",
+ timeout=15,
+ )
+
+test_step(True, "chained five-mode network mutest completed")
diff --git a/tests/mutests/mgre_chained_clouds/u0/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/u0/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_chained_clouds/u0/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/u0/etc.frr/frr.conf
new file mode 100644
index 0000000..1aa1366
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u0/etc.frr/frr.conf
@@ -0,0 +1,17 @@
+log file /var/log/frr/frr.log
+!
+! Underlay for cloud0 (static NBMA mGRE). Routes only within cloud0's
+! own addressing; no route to any other segment. Plays no active
+! nrlsmf role -- cloud0 uses unicast peer-resolution (a static
+! neighbor table), not multicast fan-out, so no underlay multicast
+! relay is needed here.
+!
+interface eth0
+ ip address 10.0.0.1/24
+!
+interface eth1
+ ip address 10.0.1.1/24
+!
+interface eth2
+ ip address 10.0.2.1/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/u0/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/u0/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u0/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_chained_clouds/u1/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/u1/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_chained_clouds/u1/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/u1/etc.frr/frr.conf
new file mode 100644
index 0000000..3d33df2
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u1/etc.frr/frr.conf
@@ -0,0 +1,16 @@
+log file /var/log/frr/frr.log
+!
+! Underlay for cloud1 (NHRP-resolved mGRE). Disconnected from every
+! other segment at the IP layer, same as u0. No nrlsmf role: NHRP
+! resolves peers to unicast underlay addresses, so no multicast
+! relay is needed on this underlay either.
+!
+interface eth0
+ ip address 10.1.0.1/24
+!
+interface eth1
+ ip address 10.1.1.1/24
+!
+interface eth2
+ ip address 10.1.2.1/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/u1/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/u1/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u1/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_chained_clouds/u2/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/u2/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_chained_clouds/u2/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/u2/etc.frr/frr.conf
new file mode 100644
index 0000000..479e80a
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u2/etc.frr/frr.conf
@@ -0,0 +1,19 @@
+log file /var/log/frr/frr.log
+!
+! Underlay for cloud2 (multicast-underlay mGRE). Unlike u0/u1, this
+! underlay DOES run nrlsmf -- `rmerge` across all three of its LANs,
+! as a stand-in for real multicast routing (e.g. PIM) -- since
+! cloud2's GRE tunnels rely on underlay multicast fan-out rather than
+! a per-peer table. This is the one place in this topology where
+! nrlsmf uses `merge`/`rmerge` instead of `cf`: it's acting as the
+! underlay's multicast routing, not as an overlay SMF router.
+!
+interface eth0
+ ip address 10.2.0.1/24
+!
+interface eth1
+ ip address 10.2.1.1/24
+!
+interface eth2
+ ip address 10.2.2.1/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/u2/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/u2/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u2/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_chained_clouds/u3/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/u3/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_chained_clouds/u3/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/u3/etc.frr/frr.conf
new file mode 100644
index 0000000..28b1c48
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u3/etc.frr/frr.conf
@@ -0,0 +1,12 @@
+log file /var/log/frr/frr.log
+!
+! Underlay for cloud3 (point-to-point GRE). Two peers, one fixed
+! remote each way. Disconnected from every other segment; no nrlsmf
+! role needed.
+!
+interface eth0
+ ip address 10.3.0.1/24
+!
+interface eth1
+ ip address 10.3.1.1/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/u3/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/u3/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u3/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_chained_clouds/u4/etc.frr/daemons b/tests/mutests/mgre_chained_clouds/u4/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_chained_clouds/u4/etc.frr/frr.conf b/tests/mutests/mgre_chained_clouds/u4/etc.frr/frr.conf
new file mode 100644
index 0000000..6b35f26
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u4/etc.frr/frr.conf
@@ -0,0 +1,16 @@
+log file /var/log/frr/frr.log
+!
+! Underlay for cloud4 (external/metadata GRE). Disconnected from
+! every other segment. No nrlsmf role: encapsulation is resolved by
+! per-destination lwtunnel routes on each peer, not by anything the
+! underlay itself needs to do.
+!
+interface eth0
+ ip address 10.4.0.1/24
+!
+interface eth1
+ ip address 10.4.1.1/24
+!
+interface eth2
+ ip address 10.4.2.1/24
+!
diff --git a/tests/mutests/mgre_chained_clouds/u4/etc.frr/vtysh.conf b/tests/mutests/mgre_chained_clouds/u4/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_chained_clouds/u4/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_four_peers/README.md b/tests/mutests/mgre_four_peers/README.md
index 4ddba07..2ad3997 100644
--- a/tests/mutests/mgre_four_peers/README.md
+++ b/tests/mutests/mgre_four_peers/README.md
@@ -1,29 +1,71 @@
-# mGRE four-peer mutests
+# GRE / mGRE four-router mutests
-Hub-and-spoke underlay with GRE overlay among four peers.
+Four routers (r0..r3) building GRE/mGRE overlay tunnels across a shared
+underlay router (u0).
```
-p1 -- lan1 --\
-p2 -- lan2 ---\
- r0 (static/connected routes)
-p3 -- lan3 ---/
-p4 -- lan4 --/
+ h0 -- r0 -- lan0 --\
+ h1 -- r1 -- lan1 ---\
+ u0 (underlay: routes/relays between the LANs)
+ h2 -- r2 -- lan2 ---/
+ h3 -- r3 -- lan3 --/
```
+`h0`..`h3` are application hosts off `r0`..`r3` (iperf source/receivers; not GRE/SMF nodes).
+Each SMF router CFs its host LAN (`eth1`) plus its GRE iface so nrlsmf is
+both the first hop onto the overlay and the last hop off it. Overlay GRE
+ifaces are `layered` so a packet received on the tunnel is not flooded
+back out of it, except the NHRP hub (`r3`): nhrpd only programs
+spoke→hub and hub→spokes, so the hub must replicate overlay multicast.
+
+`u0` is the "underlay" node — it plays the role of the WAN/network that
+a real GRE tunnel would cross, and that the operator running the
+overlay generally does *not* control. It runs no GRE/nrlsmf overlay of
+its own except in the multicast-underlay test, where it stands in for
+real underlay multicast routing (PIM) as a pure dataplane relay. It is
+never given an NHRP role, for the same reason: NHRP is part of the
+overlay's own control plane and belongs on operator-controlled
+infrastructure, not on the untrusted transit network.
+
+`r0`..`r3` are the routers that actually build tunnels and run nrlsmf.
+In the NHRP test, `r3` additionally acts as the hub / NHRP Server —
+still one of the operator's own four routers, just doing double duty.
+
## Tests
-| File | Example use case |
-|------|------------------|
-| `mutest_mgre.py` | Multipoint GRE with NBMA neighbors (unicast underlay through `r0`); peer `nrlsmf` CF with `map`/`ujoin` |
-| `mutest_mgre_mcast.py` | Multicast-remote GRE (`mgre0`); `nrlsmf rmerge` on `r0` floods the underlay group; peer CF + overlay multicast |
+| File | Mode | What resolves "which peer?" |
+|------|------|------------------------------|
+| `mutest_gre_p2p.py` | Point-to-point GRE (two independent pairs: r0<->r1, r2<->r3) | N/A — each tunnel has exactly one fixed peer |
+| `mutest_mgre_static.py` | Multipoint GRE, static NBMA | Kernel `ip neigh` for overlay unicast; nrlsmf `map ,,dynamic` learns those dests for overlay multicast inject. Ends with runtime `with-frr` + `elastic overlay` on the same `eth1,gre1` group (keep h1, stop h2/h3; idle hosts at most 1 pps). |
+| `mutest_mgre_nhrp.py` | Multipoint GRE, NHRP-resolved | FRR `nhrpd` programs kernel `ip neigh` (hub/NHS on `r3`, never `u0`); nrlsmf `map …,dynamic` learns those dests. Spokes are layered; the hub is not (overlay-mcast replicator). |
+| `mutest_mgre_mcast.py` | Multipoint GRE, multicast underlay remote | Nothing — underlay multicast fan-out delivers to every peer from one transmission |
+| `mutest_gre_external.py` | "External" (metadata) GRE | Per-destination lwtunnel routes for overlay unicast; nrlsmf overlay-multicast inject uses explicit `map ,,` (`dynamic` does not apply — no `ip neigh`). Skipped on Linux < 5.0. |
+
+Each file's docstring explains its mode in detail, including how it
+compares to the others and what it demonstrates about nrlsmf's `map`/
+`ujoin`/`uleave` commands. Static NBMA and NHRP overlay multicast use
+`map ,,dynamic` (learn from `ip neigh`; nhrpd fills that
+table in the NHRP test, and the hub is the overlay-mcast replicator). Multicast-underlay mGRE uses `ujoin`.
+External/metadata GRE uses explicit unicast `map`s because the device
+has no endpoints to auto-discover and no `ip neigh` table to learn.
+`0.0.0.0` is the kernel wildcard remote, not a learn switch.
+
+None of these tests put the overlay on the kernel's fallback `gre0`
+(remote any / local any). That device steals inbound GRE and the
+overlay subnet if it is left addressed or up. Every variant creates a
+dedicated tunnel (`gre1`, or `mgre0` for the multicast-remote case)
+and flushes/`down`s `gre0` first.
## Run
-From `tests/mutests` (requires root, FRR, `nrlsmf` on PATH):
+From `tests/mutests` (requires root, FRR with `nhrpd` enabled, `nrlsmf` on PATH):
```bash
sudo mutest mgre_four_peers
# or one file:
-sudo mutest mgre_four_peers/mutest_mgre.py
+sudo mutest mgre_four_peers/mutest_gre_p2p.py
+sudo mutest mgre_four_peers/mutest_mgre_static.py
+sudo mutest mgre_four_peers/mutest_mgre_nhrp.py
sudo mutest mgre_four_peers/mutest_mgre_mcast.py
+sudo mutest mgre_four_peers/mutest_gre_external.py
```
diff --git a/tests/mutests/mgre_four_peers/four_peer_hosts.py b/tests/mutests/mgre_four_peers/four_peer_hosts.py
new file mode 100644
index 0000000..861b81b
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/four_peer_hosts.py
@@ -0,0 +1,164 @@
+"""Shared host-LAN helpers for the four-peer GRE/mGRE mutests.
+
+Each router has an application host off eth1. Iperf is sourced on h0
+(off r0) and received on h1/h2/h3 (off r1/r2/r3) so nrlsmf CF is both
+the first hop onto the overlay and the last hop off it.
+"""
+
+import time
+
+HOST_IFACE = "eth0"
+ROUTER_HOST_IFACE = "eth1"
+IPERF_TTL = "16"
+
+# (router, router eth1 addr, host name, host addr)
+HOST_LANS = (
+ ("r0", "192.168.55.1", "h0", "192.168.55.2"),
+ ("r1", "192.168.56.1", "h1", "192.168.56.2"),
+ ("r2", "192.168.57.1", "h2", "192.168.57.2"),
+ ("r3", "192.168.58.1", "h3", "192.168.58.2"),
+)
+
+SOURCE_HOST = "h0"
+SOURCE_HOST_ADDR = "192.168.55.2"
+RECV_HOSTS = ("h1", "h2", "h3")
+
+# Back-compat names used by older call sites
+HOST = SOURCE_HOST
+HOST_ADDR = SOURCE_HOST_ADDR
+R0_HOST_IFACE = ROUTER_HOST_IFACE
+R0_HOST_ADDR = "192.168.55.1"
+
+
+def setup_host_lan(step, wait_step):
+ """Bring up each router's host LAN and its application host."""
+ for router, router_addr, host, host_addr in HOST_LANS:
+ step(router, f"ethtool -K {ROUTER_HOST_IFACE} rx off tx off || true")
+ wait_step(
+ router,
+ f"ip -br addr show dev {ROUTER_HOST_IFACE}",
+ match=router_addr,
+ desc=f"{router} {ROUTER_HOST_IFACE} address {router_addr}",
+ timeout=30,
+ )
+ step(router, "sysctl -w net.ipv4.conf.all.mc_forwarding=0 || true")
+
+ step(host, f"ethtool -K {HOST_IFACE} rx off tx off || true")
+ step(host, f"ip addr add {host_addr}/24 dev {HOST_IFACE} || true")
+ step(host, f"ip link set {HOST_IFACE} up")
+ step(host, f"ip route replace 0.0.0.0/0 via {router_addr}")
+ wait_step(
+ host,
+ f"ip -br addr show dev {HOST_IFACE}",
+ match=host_addr,
+ desc=f"{host} {HOST_IFACE} address {host_addr}",
+ timeout=30,
+ )
+ wait_step(
+ host,
+ f"ping -c1 -W2 {router_addr}",
+ match="1 received",
+ desc=f"{host} reaches {router} on the host LAN",
+ timeout=20,
+ )
+
+
+def start_overlay_mcast_servers(step, receivers, mcast, gre_dev=None):
+ for name in receivers:
+ step(name, f"ip route replace {mcast}/32 dev {HOST_IFACE}")
+ step(
+ name,
+ f"iperf -u -T 4 -i 1 -s -e -B {mcast}%{HOST_IFACE} "
+ f"> iperf-{name}-server.log 2>&1 &",
+ )
+
+
+def restart_overlay_mcast_servers(step, receivers, mcast):
+ """New iperf server log. Do not truncate a live iperf file: the
+ writer keeps its old offset (sparse NULs) and grep then treats the
+ log as binary and never prints ``8 pps``.
+ """
+ for name in receivers:
+ step(name, "pkill iperf || true")
+ time.sleep(1)
+ start_overlay_mcast_servers(step, receivers, mcast)
+
+
+def start_host_mcast_client(step, wait_step, mcast):
+ step(SOURCE_HOST, f"ip route replace {mcast}/32 dev {HOST_IFACE}")
+ step(
+ SOURCE_HOST,
+ f"iperf -u -T {IPERF_TTL} -t 1000 -i 1 -b 8pps -l 1024 -e "
+ f"-B {SOURCE_HOST_ADDR} -c {mcast} &> iperf-{SOURCE_HOST}-client.log &",
+ )
+ wait_step(
+ SOURCE_HOST,
+ f"tail -n1 iperf-{SOURCE_HOST}-client.log",
+ match="8 pps",
+ desc="h0 sending application multicast at 8 pps",
+ timeout=30,
+ )
+
+
+def wait_overlay_mcast_receivers(wait_step, receivers):
+ for name in receivers:
+ wait_step(
+ name,
+ f'grep "8 pps" iperf-{name}-server.log',
+ match="8 pps",
+ desc=f"{name} receiving application multicast at 8 pps",
+ timeout=20,
+ )
+
+
+def count_overlay_mcast_pkts(step, node, mcast, iface=HOST_IFACE, window_s=2):
+ """Count packets to ``mcast`` on ``iface`` over ``window_s`` seconds."""
+ raw = step(
+ node,
+ f"timeout {window_s} tcpdump -nn -l -i {iface} host {mcast} 2>/dev/null "
+ f"| grep -c {mcast} || true",
+ )
+ n = 0
+ for tok in str(raw).split():
+ if tok.isdigit():
+ n = int(tok)
+ return n
+
+
+def cleanup_iperf(step, receivers):
+ step(SOURCE_HOST, "pkill iperf || true")
+ for name in receivers:
+ step(name, "pkill iperf || true")
+
+
+def enable_host_igmp(step, wait_step, routers, host_iface=ROUTER_HOST_IFACE):
+ """Enable IGMP on each router's host LAN.
+
+ FRR serves IGMP from pimd, so that daemon must be running for
+ ``show ip igmp`` / nrlsmf ``with-frr``. No ``ip pim`` on the iface.
+ """
+ for name in routers:
+ step(name, "pgrep -x pimd >/dev/null || /usr/lib/frr/pimd -d")
+ step(
+ name,
+ "vtysh -c 'configure terminal' "
+ f"-c 'interface {host_iface}' "
+ "-c 'ip igmp'",
+ )
+ wait_step(
+ name,
+ f"vtysh -c 'show ip igmp interface {host_iface}'",
+ match=host_iface,
+ desc=f"{name} FRR IGMP enabled on {host_iface}",
+ timeout=20,
+ )
+
+
+def wait_igmp_group(wait_step, router, mcast, timeout=30):
+ wait_step(
+ router,
+ "vtysh -c 'show ip igmp groups'",
+ match=mcast,
+ desc=f"{router} FRR IGMP has {mcast}",
+ timeout=timeout,
+ )
diff --git a/tests/mutests/mgre_four_peers/munet.yaml b/tests/mutests/mgre_four_peers/munet.yaml
index 5d61a5d..6c1f254 100644
--- a/tests/mutests/mgre_four_peers/munet.yaml
+++ b/tests/mutests/mgre_four_peers/munet.yaml
@@ -1,40 +1,75 @@
-# Four mGRE peers with a hub router between underlay LANs.
+# Four GRE/mGRE routers with an underlay router between the LANs,
+# plus an application host off each SMF router.
#
-# p1 -- lan1 --\
-# p2 -- lan2 ---\
-# r0
-# p3 -- lan3 ---/
-# p4 -- lan4 --/
+# h0 -- r0 -- lan0 --\
+# h1 -- r1 -- lan1 ---\
+# u0
+# h2 -- r2 -- lan2 ---/
+# h3 -- r3 -- lan3 --/
#
-# Underlay: static/connected routes via r0.
-# Overlay: multipoint GRE (remote any) + NBMA neighbors; nrlsmf CF on gre0.
+# u0 is the "underlay" node: it does no GRE/nrlsmf overlay work of its
+# own (except in the multicast-underlay test, where it acts purely as a
+# multicast dataplane relay between the LANs). It exists only to route
+# ordinary IP traffic between r0..r3, standing in for "the WAN" that a
+# real GRE/mGRE deployment would tunnel across.
+#
+# r0..r3 build GRE/mGRE tunnels and run nrlsmf. Application multicast
+# is sourced on host h0 and received on h1/h2/h3 so nrlsmf CF is both
+# the first hop onto the overlay and the last hop off it — not a
+# locally generated send or receive on the tunnel.
+# Each test reuses this topology and only changes what r0..r3 (and
+# sometimes u0) do on top of it.
topology:
networks:
+ - name: lan0
- name: lan1
- name: lan2
- name: lan3
- - name: lan4
+ - name: lanh0
+ - name: lanh1
+ - name: lanh2
+ - name: lanh3
nodes:
- - name: r0
+ - name: u0
kind: frr
connections:
+ - to: lan0
- to: lan1
- to: lan2
- to: lan3
- - to: lan4
- - name: p1
+ - name: h0
+ kind: host
+ connections:
+ - to: lanh0
+ - name: h1
+ kind: host
+ connections:
+ - to: lanh1
+ - name: h2
+ kind: host
+ connections:
+ - to: lanh2
+ - name: h3
+ kind: host
+ connections:
+ - to: lanh3
+ - name: r0
+ kind: frr
+ connections:
+ - to: lan0
+ - to: lanh0
+ - name: r1
kind: frr
connections:
- to: lan1
- - name: p2
+ - to: lanh1
+ - name: r2
kind: frr
connections:
- to: lan2
- - name: p3
+ - to: lanh2
+ - name: r3
kind: frr
connections:
- to: lan3
- - name: p4
- kind: frr
- connections:
- - to: lan4
+ - to: lanh3
diff --git a/tests/mutests/mgre_four_peers/mutest_gre_external.py b/tests/mutests/mgre_four_peers/mutest_gre_external.py
new file mode 100644
index 0000000..b25b0f2
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/mutest_gre_external.py
@@ -0,0 +1,364 @@
+"""Example: "external" (metadata) GRE + nrlsmf CF among four routers.
+
+Topology (shared with other tests in this directory):
+
+ h0 -- r0 -- lan0 --\\
+ r1 -- lan1 ---\\
+ u0 (underlay: routes ordinary IP between the LANs)
+ r2 -- lan2 ---/
+ r3 -- lan3 --/
+
+What "external" (metadata) GRE means here
+--------------------------------------------
+Every other GRE/mGRE mode in this directory fixes its encapsulation
+parameters (local address, and either a fixed remote or a wildcard
+remote) directly *on the tunnel interface* -- that's what lets nrlsmf
+read them straight back out via netlink with no configuration of its
+own. An "external" (a.k.a. "collect metadata" or "lightweight") GRE
+device has none of that: `ip link add ... type gre external` creates
+an interface with no fixed local/remote/key at all. Instead, whatever
+adds routes or flow rules for it supplies the encapsulation parameters
+per destination, using Linux's lightweight tunnel ("lwtunnel") route
+encap:
+
+ ip route add /32 encap ip id src dst \\
+ ttl dev gre1
+
+This is the mechanism SDN controllers and OVS/OVN typically use to
+build many per-flow or per-peer tunnels dynamically out of a single
+device, rather than the operator hand-configuring one interface per
+peer. It's conceptually the multipoint-resolution equivalent of the
+static NBMA table in mutest_mgre_static.py (mutest_mgre_static.py
+resolves peers via `ip neigh` entries; this resolves them via
+per-destination routes instead) -- just with the resolution table
+expressed as routes rather than neighbor entries, and populated
+externally rather than beingsomething nrlsmf or the kernel's GRE
+driver can discover on its own.
+
+Why this is the one case where nrlsmf's `map` command is required
+---------------------------------------------------------------------
+nrlsmf normally reads a GRE interface's local/remote endpoint addresses
+straight from the kernel and needs no `map` command at all -- true for
+every mode in this directory except this one. An external GRE device
+reports no fixed local or remote address to read (they don't exist on
+the interface itself), so nrlsmf has nothing to auto-discover. This
+test deliberately starts nrlsmf once *without* `map` to show the
+warning nrlsmf logs in that situation, then again *with* explicit
+per-peer `map gre1,,` entries. Overlay unicast still
+uses the kernel lwtunnel routes (independent of `map`). Overlay
+multicast inject onto this device has no single remote, so nrlsmf
+transmits once per mapped unicast peer -- same send path as static
+NBMA / NHRP. `map …,0.0.0.0` only records the wildcard; it is not a
+send dest. `map …,dynamic` does not apply here (no `ip neigh` table).
+
+What this example covers
+-------------------------
+* Underlay: unicast IP between all four routers, routed through u0.
+* Overlay: one external GRE device (gre1) per router, sharing a single
+ 172.16.0.0/24 overlay subnet, with per-destination lwtunnel routes
+ providing the encapsulation parameters to reach each other router --
+ the "external" analogue of the NBMA table in mutest_mgre_static.py. gre1
+ is used because the kernel's built-in fallback gre0 cannot be turned
+ into a collect-md device (same reason mutest_mgre_mcast.py uses mgre0).
+* nrlsmf classic flooding (`cf`) on each router's host LAN plus gre1:
+ - First started *without* `map`, to confirm nrlsmf logs the
+ documented "missing tunnel endpoint addressing... must map it"
+ warning. Overlay unicast still works (kernel lwtunnel routes).
+ - Then restarted with explicit `map gre1,,` for every
+ other router, so overlay-multicast inject has unicast remotes.
+ - Iperf sourced on host h0, received at h1/h2/h3.
+
+See mutest_gre_p2p.py (point-to-point), mutest_mgre_static.py (static
+NBMA mGRE), mutest_mgre_nhrp.py (NHRP-resolved mGRE), and
+mutest_mgre_mcast.py (multicast-underlay mGRE) for the other GRE
+tunnel modes.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from four_peer_hosts import RECV_HOSTS
+from four_peer_hosts import cleanup_iperf
+from four_peer_hosts import setup_host_lan
+from four_peer_hosts import start_host_mcast_client
+from four_peer_hosts import start_overlay_mcast_servers
+from four_peer_hosts import wait_overlay_mcast_receivers
+from kernel_compat import min_kernel_version
+from smf_cli import check_common_show
+from smf_cli import check_show_neighbors
+from smf_cli import check_show_tunnel
+from smf_cli import show_json
+
+ROUTERS = {
+ "r0": {"underlay": "10.0.0.2", "overlay": "172.16.0.1"},
+ "r1": {"underlay": "10.0.1.2", "overlay": "172.16.0.2"},
+ "r2": {"underlay": "10.0.2.2", "overlay": "172.16.0.3"},
+ "r3": {"underlay": "10.0.3.2", "overlay": "172.16.0.4"},
+}
+
+# Dedicated name: kernel fallback gre0 is not a collect-md device and
+# will steal inbound GRE (and the overlay subnet) if we try to reuse it.
+GRE_DEV = "gre1"
+GRE_KEY = "100"
+MISSING_ENDPOINT_WARNING = "missing tunnel endpoint addressing"
+OVERLAY_MCAST = "239.0.0.1"
+
+if min_kernel_version((5, 0)):
+ return "skip"
+
+
+section("Disable offloads and wait for underlay addresses")
+
+step("u0", "sysctl -w net.ipv4.ip_forward=1")
+
+for name, cfg in ROUTERS.items():
+ step(name, "ethtool -K eth0 rx off tx off || true")
+ wait_step(
+ name,
+ "ip -br addr show dev eth0",
+ match=cfg["underlay"],
+ desc=f"{name} underlay address {cfg['underlay']}",
+ timeout=30,
+ )
+
+for ifname, addr in (
+ ("eth0", "10.0.0.1"),
+ ("eth1", "10.0.1.1"),
+ ("eth2", "10.0.2.1"),
+ ("eth3", "10.0.3.1"),
+):
+ step("u0", f"ethtool -K {ifname} rx off tx off || true")
+ wait_step(
+ "u0",
+ f"ip -br addr show dev {ifname}",
+ match=addr,
+ desc=f"u0 {ifname} address {addr}",
+ timeout=30,
+ )
+
+setup_host_lan(step, wait_step)
+
+section("Underlay unicast reachability through u0")
+
+for src in ROUTERS:
+ for dst, dcfg in ROUTERS.items():
+ if src == dst:
+ continue
+ wait_step(
+ src,
+ f"ping -c1 -W2 {dcfg['underlay']}",
+ match="1 received",
+ desc=f"{src} ping underlay {dst} ({dcfg['underlay']})",
+ timeout=20,
+ )
+
+section("Create external (metadata) GRE devices with per-destination routes")
+
+for name, cfg in ROUTERS.items():
+ # Don't reuse kernel fallback gre0 as the external device: that
+ # `ip link add` often fails (device exists), step() keeps going,
+ # and overlay pings then hit a remote-any tunnel with no remotes.
+ step(name, "ip addr flush dev gre0 2>/dev/null || true")
+ step(name, "ip link set gre0 down 2>/dev/null || true")
+ step(name, f"ip link del {GRE_DEV} 2>/dev/null || true")
+ # No local/remote/key here at all -- that's what makes this
+ # "external": the device carries none of its own encapsulation
+ # parameters.
+ step(name, f"ip link add name {GRE_DEV} type gre external")
+ step(name, f"ip link set {GRE_DEV} multicast on")
+ step(name, f"ip link set {GRE_DEV} up")
+ step(name, f"ip addr add {cfg['overlay']}/24 dev {GRE_DEV}")
+ # Per-destination lwtunnel routes supply the encapsulation
+ # parameters that would otherwise live on the interface. This is
+ # the "external" analogue of the ip-neigh NBMA table in
+ # mutest_mgre_static.py -- same job, different mechanism.
+ for other, ocfg in ROUTERS.items():
+ if other == name:
+ continue
+ step(
+ name,
+ f"ip route replace {ocfg['overlay']}/32 encap ip id {GRE_KEY} "
+ f"src {cfg['underlay']} dst {ocfg['underlay']} ttl 64 "
+ f"dev {GRE_DEV}",
+ )
+ wait_step(
+ name,
+ f"ip -br link show {GRE_DEV}",
+ match="UP",
+ desc=f"{name} {GRE_DEV} is UP",
+ )
+ wait_step(
+ name,
+ f"ip -d link show {GRE_DEV}",
+ match="external",
+ desc=f"{name} {GRE_DEV} is collect-md / external",
+ )
+ # Confirm one lwtunnel route actually installed (not a plain
+ # on-link /24 leftover from a failed encap command).
+ peer = next(n for n in ROUTERS if n != name)
+ wait_step(
+ name,
+ f"ip route get {ROUTERS[peer]['overlay']}",
+ match="encap",
+ desc=f"{name} overlay route to {peer} uses lwtunnel encap",
+ )
+
+section("Overlay unicast across external GRE (kernel-level, before nrlsmf)")
+
+# This works purely from the lwtunnel routes above -- nothing here
+# depends on nrlsmf or its map state, which is the point being made in
+# the next section.
+for src, scfg in ROUTERS.items():
+ for dst, dcfg in ROUTERS.items():
+ if src == dst:
+ continue
+ wait_step(
+ src,
+ f"ping -c1 -W3 -I {scfg['overlay']} {dcfg['overlay']}",
+ match="1 received",
+ desc=f"{src} ping overlay {dst} ({dcfg['overlay']})",
+ timeout=20,
+ )
+
+section("Start nrlsmf WITHOUT map -- confirm the documented warning fires")
+
+for name in ROUTERS:
+ step(
+ name,
+ "nrlsmf debug 4 "
+ f"instance smf-{name}-ext-nomap "
+ f"add overlay,cf,eth1,{GRE_DEV} "
+ f"layered {GRE_DEV} "
+ "&> nrlsmf-ext-nomap.log &",
+ )
+ wait_step(
+ name,
+ f'pgrep -af "nrlsmf.*instance smf-{name}-ext-nomap"',
+ match=f"smf-{name}-ext-nomap",
+ desc=f"{name} nrlsmf running on {GRE_DEV} (no map)",
+ timeout=20,
+ )
+ wait_step(
+ name,
+ "grep "
+ f'"{MISSING_ENDPOINT_WARNING}" nrlsmf-ext-nomap.log',
+ match=GRE_DEV,
+ desc=f"{name} nrlsmf logs the missing-endpoint warning for {GRE_DEV}",
+ timeout=20,
+ )
+
+section("Overlay unicast still works without map (kernel handles it, not nrlsmf)")
+
+wait_step(
+ "r0",
+ f"ping -c1 -W3 -I {ROUTERS['r0']['overlay']} {ROUTERS['r1']['overlay']}",
+ match="1 received",
+ desc="r0 ping overlay r1 (still fine -- map affects nrlsmf bookkeeping only)",
+ timeout=20,
+)
+
+section("nrlsmf --cli show tunnel (json, no map -- no configured remotes)")
+
+nomap = show_json("r0", "show tunnel", "smf-r0-ext-nomap")
+if nomap is not None:
+ test_step(isinstance(nomap, list), "r0 nomap show tunnel json is a list", target="r0")
+ gre_rows = [r for r in nomap if isinstance(r, dict) and r.get("Interface") == GRE_DEV]
+ test_step(
+ not any("C" in (r.get("Flags") or "") for r in gre_rows),
+ "r0 nomap show tunnel has no C flag on gre1",
+ target="r0",
+ )
+
+section("Stop the un-mapped instances")
+
+for name in ROUTERS:
+ step(name, "pkill nrlsmf || true")
+ wait_step(
+ name,
+ "pgrep -af nrlsmf || true",
+ match="",
+ desc=f"{name} nrlsmf stopped",
+ timeout=15,
+ )
+
+section("Restart nrlsmf WITH per-peer map -- overlay mcast inject dests")
+
+for name, cfg in ROUTERS.items():
+ maps = " ".join(
+ f"map {GRE_DEV},{cfg['underlay']},{ocfg['underlay']}"
+ for other, ocfg in ROUTERS.items()
+ if other != name
+ )
+ step(
+ name,
+ "nrlsmf debug 4 "
+ f"instance smf-{name}-ext "
+ f"add overlay,cf,eth1,{GRE_DEV} "
+ f"layered {GRE_DEV} "
+ f"{maps} "
+ "&> nrlsmf-ext.log &",
+ )
+ wait_step(
+ name,
+ f'pgrep -af "nrlsmf.*instance smf-{name}-ext"',
+ match=f"smf-{name}-ext",
+ desc=f"{name} nrlsmf running on {GRE_DEV} (mapped)",
+ timeout=20,
+ )
+ wait_step(
+ name,
+ 'grep "regular group" nrlsmf-ext.log',
+ match="overlay",
+ desc=f"{name} nrlsmf log shows overlay group",
+ timeout=20,
+ )
+
+section("nrlsmf --cli show tunnel / neighbors (json, explicit map)")
+
+for name, cfg in ROUTERS.items():
+ peers = [ocfg for other, ocfg in ROUTERS.items() if other != name]
+ inst = f"smf-{name}-ext"
+ check_common_show(name, inst, group_name="overlay", ifaces=("eth1", GRE_DEV))
+ check_show_tunnel(
+ name, inst, GRE_DEV,
+ local=cfg["underlay"],
+ remotes=[p["underlay"] for p in peers],
+ overlay_ip=cfg["overlay"],
+ want_c=True,
+ )
+ # External GRE has no ip neigh table; mapped remotes appear as Neighbor IP "-".
+ check_show_neighbors(
+ name, inst, GRE_DEV,
+ remotes=[p["underlay"] for p in peers],
+ min_count=3,
+ want_c=True,
+ )
+
+section("[External] Overlay multicast: h0 -> SMF -> h1/h2/h3")
+
+RECEIVERS = RECV_HOSTS
+start_overlay_mcast_servers(step, RECEIVERS, OVERLAY_MCAST)
+start_host_mcast_client(step, wait_step, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, RECEIVERS)
+
+section("Cleanup")
+
+cleanup_iperf(step, RECEIVERS)
+for name in ROUTERS:
+ step(name, "pkill nrlsmf || true")
+ wait_step(
+ name,
+ "pgrep -af nrlsmf || true",
+ match="",
+ desc=f"{name} nrlsmf stopped",
+ timeout=15,
+ )
+
+test_step(True, "External (metadata) GRE four-router mutest completed")
diff --git a/tests/mutests/mgre_four_peers/mutest_gre_p2p.py b/tests/mutests/mgre_four_peers/mutest_gre_p2p.py
new file mode 100644
index 0000000..2252310
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/mutest_gre_p2p.py
@@ -0,0 +1,262 @@
+"""Example: point-to-point GRE + nrlsmf CF between paired routers.
+
+Topology (shared with other tests in this directory):
+
+ h0 -- r0 -- lan0 --\\
+ r1 -- lan1 ---\\
+ u0 (underlay: routes ordinary IP between the LANs)
+ r2 -- lan2 ---/
+ r3 -- lan3 --/
+
+What "point-to-point GRE" means here
+-------------------------------------
+A P2P GRE tunnel has exactly one local address and one remote address,
+both fixed when the tunnel is created. There's no question of "which
+peer does this packet go to" -- there's only ever one peer -- so unlike
+the multipoint (mGRE) modes in the other files in this directory, no
+peer-resolution mechanism (static table or NHRP) is needed at all.
+
+This test builds two independent P2P tunnels out of the four available
+routers -- r0<->r1 and r2<->r3 -- each riding over ordinary IP routing
+through u0, the same way a real P2P GRE tunnel would ride over a WAN
+hop. u0 itself is not GRE/nrlsmf-aware in this test; it's just an IP
+router in the middle.
+
+What this example covers
+-------------------------
+* Underlay: plain unicast IP reachability between the paired routers,
+ routed through u0. No direct r0<->r1 (or r2<->r3) LAN is needed --
+ the tunnel just needs IP reachability, however many hops that takes.
+* Overlay: two independent point-to-point GRE tunnels, each with its
+ own small /30 subnet so the two pairs' overlay addressing doesn't
+ collide (r0/r1 use 172.16.0.0/30, r2/r3 use 172.16.1.0/30).
+* nrlsmf on each router running classic flooding (`cf`) on its host
+ LAN plus GRE iface. Iperf is sourced on h0 and received on h1 so
+ SMF is both first hop onto and last hop off the r0<->r1 overlay.
+ No `map` command is used: endpoint addressing is auto-discovered
+ by nrlsmf from the kernel for an ordinary GRE interface.
+
+See mutest_mgre_static.py (static NBMA mGRE), mutest_mgre_nhrp.py (NHRP-resolved
+mGRE), mutest_mgre_mcast.py (multicast-underlay mGRE), and
+mutest_gre_external.py (external/metadata GRE) for the multipoint
+variants, where -- unlike here -- resolving "which peer" is the whole
+point.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from four_peer_hosts import cleanup_iperf
+from four_peer_hosts import setup_host_lan
+from four_peer_hosts import start_host_mcast_client
+from four_peer_hosts import start_overlay_mcast_servers
+from four_peer_hosts import wait_overlay_mcast_receivers
+from smf_cli import check_common_show
+from smf_cli import check_show_neighbors
+from smf_cli import check_show_tunnel
+
+# Two independent P2P GRE pairs, sharing the four-router underlay
+# topology. Each pair gets its own /30 overlay subnet so the two
+# tunnels' addressing can't be confused with each other.
+PAIRS = [
+ ("r0", "r1"),
+ ("r2", "r3"),
+]
+
+# Underlay (physical, routed-through-u0) addresses -- one /24 per LAN.
+UNDERLAY = {
+ "r0": "10.0.0.2",
+ "r1": "10.0.1.2",
+ "r2": "10.0.2.2",
+ "r3": "10.0.3.2",
+}
+
+# Overlay (GRE tunnel) addresses -- one /30 per pair:
+# r0 <-> r1 : 172.16.0.0/30 (r0=.1, r1=.2)
+# r2 <-> r3 : 172.16.1.0/30 (r2=.1, r3=.2)
+OVERLAY = {
+ "r0": "172.16.0.1",
+ "r1": "172.16.0.2",
+ "r2": "172.16.1.1",
+ "r3": "172.16.1.2",
+}
+
+# Dedicated name: kernel fallback gre0 (remote any) will steal inbound
+# GRE and the overlay subnet if we try to reuse it.
+GRE_DEV = "gre1"
+OVERLAY_MCAST = "239.0.0.1"
+
+
+section("Disable offloads and wait for underlay addresses")
+
+for name in UNDERLAY:
+ step(name, "ethtool -K eth0 rx off tx off || true")
+ wait_step(
+ name,
+ "ip -br addr show dev eth0",
+ match=UNDERLAY[name],
+ desc=f"{name} underlay address {UNDERLAY[name]}",
+ timeout=30,
+ )
+
+for ifname, addr in (
+ ("eth0", "10.0.0.1"),
+ ("eth1", "10.0.1.1"),
+ ("eth2", "10.0.2.1"),
+ ("eth3", "10.0.3.1"),
+):
+ step("u0", f"ethtool -K {ifname} rx off tx off || true")
+ wait_step(
+ "u0",
+ f"ip -br addr show dev {ifname}",
+ match=addr,
+ desc=f"u0 {ifname} address {addr}",
+ timeout=30,
+ )
+
+setup_host_lan(step, wait_step)
+
+step("u0", "sysctl -w net.ipv4.ip_forward=1")
+
+section("Underlay unicast reachability through u0")
+
+for a, b in PAIRS:
+ wait_step(
+ a,
+ f"ping -c1 -W2 {UNDERLAY[b]}",
+ match="1 received",
+ desc=f"{a} ping underlay {b} ({UNDERLAY[b]})",
+ timeout=20,
+ )
+ wait_step(
+ b,
+ f"ping -c1 -W2 {UNDERLAY[a]}",
+ match="1 received",
+ desc=f"{b} ping underlay {a} ({UNDERLAY[a]})",
+ timeout=20,
+ )
+
+section("Create point-to-point GRE tunnels")
+
+for a, b in PAIRS:
+ # Each side of the pair fixes both endpoints explicitly (local and
+ # remote). There's no wildcard/"any" remote anywhere in a P2P
+ # tunnel -- that's what distinguishes it from the mGRE cases.
+ for local, remote in ((a, b), (b, a)):
+ step(local, "ip addr flush dev gre0 2>/dev/null || true")
+ step(local, "ip link set gre0 down 2>/dev/null || true")
+ step(local, f"ip link del {GRE_DEV} 2>/dev/null || true")
+ step(
+ local,
+ f"ip link add name {GRE_DEV} type gre "
+ f"local {UNDERLAY[local]} remote {UNDERLAY[remote]} ttl 64",
+ )
+ step(local, f"ip addr add {OVERLAY[local]}/30 dev {GRE_DEV}")
+ step(local, f"ip link set {GRE_DEV} multicast on")
+ step(local, f"ip link set {GRE_DEV} up")
+ wait_step(
+ local,
+ f"ip -br link show {GRE_DEV}",
+ match="UP",
+ desc=f"{local} {GRE_DEV} is UP",
+ )
+
+section("Overlay unicast across P2P GRE (before nrlsmf)")
+
+for a, b in PAIRS:
+ wait_step(
+ a,
+ f"ping -c1 -W3 {OVERLAY[b]}",
+ match="1 received",
+ desc=f"{a} ping overlay {b} ({OVERLAY[b]})",
+ timeout=20,
+ )
+ wait_step(
+ b,
+ f"ping -c1 -W3 {OVERLAY[a]}",
+ match="1 received",
+ desc=f"{b} ping overlay {a} ({OVERLAY[a]})",
+ timeout=20,
+ )
+
+section("Start nrlsmf classic flooding on P2P GRE")
+
+for name in UNDERLAY:
+ step(
+ name,
+ "nrlsmf debug 4 "
+ f"instance smf-{name}-p2p "
+ f"add overlay,cf,eth1,{GRE_DEV} "
+ f"layered {GRE_DEV} "
+ "&> nrlsmf-gre-p2p.log &",
+ )
+ wait_step(
+ name,
+ f'pgrep -af "nrlsmf.*instance smf-{name}-p2p"',
+ match=f"smf-{name}-p2p",
+ desc=f"{name} nrlsmf running on {GRE_DEV}",
+ timeout=20,
+ )
+ wait_step(
+ name,
+ 'grep "regular group" nrlsmf-gre-p2p.log',
+ match="overlay",
+ desc=f"{name} nrlsmf log shows overlay group",
+ timeout=20,
+ )
+
+section("nrlsmf --cli show tunnel / neighbors (json, kernel-learned P2P)")
+
+# No map: Local/Remote come from the GRE device. P2P GRE is NOARP, so
+# overlay pings do not install Neighbor IP; the kernel remote is listed.
+for a, b in PAIRS:
+ inst_a = f"smf-{a}-p2p"
+ inst_b = f"smf-{b}-p2p"
+ check_common_show(a, inst_a, group_name="overlay", ifaces=("eth1", GRE_DEV))
+ check_show_tunnel(
+ a, inst_a, GRE_DEV,
+ local=UNDERLAY[a], remotes=[UNDERLAY[b]], overlay_ip=OVERLAY[a], want_c=False,
+ )
+ check_show_neighbors(
+ a, inst_a, GRE_DEV,
+ remotes=[UNDERLAY[b]], min_count=1,
+ )
+ check_show_tunnel(
+ b, inst_b, GRE_DEV,
+ local=UNDERLAY[b], remotes=[UNDERLAY[a]], overlay_ip=OVERLAY[b], want_c=False,
+ )
+ check_show_neighbors(
+ b, inst_b, GRE_DEV,
+ remotes=[UNDERLAY[a]], min_count=1,
+ )
+
+# h0 is the source (off r0); h1 (off r1) is the receiver for this pair.
+# r2<->r3 stays a unicast-only overlay check.
+section("[P2P] Overlay multicast: h0 -> SMF -> h1")
+
+RECEIVERS = ("h1",)
+start_overlay_mcast_servers(step, RECEIVERS, OVERLAY_MCAST)
+start_host_mcast_client(step, wait_step, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, RECEIVERS)
+
+section("Cleanup")
+
+cleanup_iperf(step, RECEIVERS)
+for name in UNDERLAY:
+ step(name, "pkill nrlsmf || true")
+ wait_step(
+ name,
+ "pgrep -af nrlsmf || true",
+ match="",
+ desc=f"{name} nrlsmf stopped",
+ timeout=15,
+ )
+
+test_step(True, "P2P GRE paired mutest completed")
diff --git a/tests/mutests/mgre_four_peers/mutest_mgre.py b/tests/mutests/mgre_four_peers/mutest_mgre.py
deleted file mode 100644
index 0c608d2..0000000
--- a/tests/mutests/mgre_four_peers/mutest_mgre.py
+++ /dev/null
@@ -1,178 +0,0 @@
-"""Example: multipoint GRE (NBMA) + nrlsmf CF among four peers.
-
-Topology (shared with other tests in this directory):
-
- p1 -- lan1 --\
- p2 -- lan2 ---\
- r0 (hub; static/connected underlay routes)
- p3 -- lan3 ---/
- p4 -- lan4 --/
-
-What this example covers
-------------------------
-* Underlay: unicast IP between peers through hub router r0.
-* Overlay: multipoint GRE (remote any) with static NBMA neighbors
- (overlay IP -> underlay IP), similar to a static DMVPN-style mesh.
-* nrlsmf on each peer running classic flooding on gre0, with GRE helpers:
- - map gre0,,0.0.0.0 (mGRE / any-remote tunnel mapping)
- - ujoin ,eth0 (underlay group join API)
-
-See mutest_mgre_mcast.py for the underlay-multicast GRE remote example.
-"""
-
-from munet.mutest.userapi import section
-from munet.mutest.userapi import step
-from munet.mutest.userapi import test_step
-from munet.mutest.userapi import wait_step
-
-# Underlay LAN addressing (matches */etc.frr/frr.conf)
-PEERS = {
- "p1": {"underlay": "10.0.1.2", "overlay": "172.16.0.1"},
- "p2": {"underlay": "10.0.2.2", "overlay": "172.16.0.2"},
- "p3": {"underlay": "10.0.3.2", "overlay": "172.16.0.3"},
- "p4": {"underlay": "10.0.4.2", "overlay": "172.16.0.4"},
-}
-
-# Used with ujoin (GRE underlay join API).
-UNDERLAY_MCAST = "239.1.1.1"
-OVERLAY_MCAST = "239.0.0.1"
-GRE_DEV = "gre0"
-
-
-def peer_names():
- return list(PEERS.keys())
-
-
-section("Disable offloads and wait for underlay addresses")
-
-step("r0", "sysctl -w net.ipv4.ip_forward=1")
-
-for name, cfg in PEERS.items():
- step(name, "ethtool -K eth0 rx off tx off || true")
- wait_step(
- name,
- "ip -br addr show dev eth0",
- match=cfg["underlay"],
- desc=f"{name} underlay address {cfg['underlay']}",
- timeout=30,
- )
-
-for ifname, addr in (
- ("eth0", "10.0.1.1"),
- ("eth1", "10.0.2.1"),
- ("eth2", "10.0.3.1"),
- ("eth3", "10.0.4.1"),
-):
- step("r0", f"ethtool -K {ifname} rx off tx off || true")
- wait_step(
- "r0",
- f"ip -br addr show dev {ifname}",
- match=addr,
- desc=f"r0 {ifname} address {addr}",
- timeout=30,
- )
-
-section("Underlay unicast reachability through hub router")
-
-for src in peer_names():
- for dst, dcfg in PEERS.items():
- if src == dst:
- continue
- wait_step(
- src,
- f"ping -c1 -W2 {dcfg['underlay']}",
- match="1 received",
- desc=f"{src} ping underlay {dst} ({dcfg['underlay']})",
- timeout=20,
- )
-
-section("Create multipoint GRE tunnels with NBMA neighbors")
-
-for name, cfg in PEERS.items():
- step(name, f"ip link del {GRE_DEV} 2>/dev/null || true")
- step(
- name,
- f"ip link add name {GRE_DEV} type gre "
- f"local {cfg['underlay']} remote 0.0.0.0 ttl 64",
- )
- step(name, f"ip addr add {cfg['overlay']}/24 dev {GRE_DEV}")
- step(name, f"ip link set {GRE_DEV} up")
- # Map each remote overlay address to that peer's underlay endpoint.
- for other, ocfg in PEERS.items():
- if other == name:
- continue
- step(
- name,
- f"ip neigh replace {ocfg['overlay']} lladdr {ocfg['underlay']} "
- f"nud permanent dev {GRE_DEV}",
- )
- step(name, f"ip route replace {OVERLAY_MCAST}/32 dev {GRE_DEV}")
- wait_step(
- name,
- f"ip -br link show {GRE_DEV}",
- match="UP",
- desc=f"{name} {GRE_DEV} is UP",
- )
-
-section("Overlay unicast across mGRE (before nrlsmf)")
-
-for src, scfg in PEERS.items():
- for dst, dcfg in PEERS.items():
- if src == dst:
- continue
- wait_step(
- src,
- f"ping -c1 -W3 -I {scfg['overlay']} {dcfg['overlay']}",
- match="1 received",
- desc=f"{src} ping overlay {dst} ({dcfg['overlay']})",
- timeout=20,
- )
-
-section("Start nrlsmf classic flooding on mGRE (map + ujoin)")
-
-for name, cfg in PEERS.items():
- # nrlsmf startup for GRE/mGRE:
- # instance smf-{name} — unique control pipe (/tmp is shared across munet ns)
- # add overlay,cf,gre0 — classic flooding on the GRE iface
- # map gre0,local,0.0.0.0 — GRE: tunnel endpoints; 0.0.0.0 = mGRE/any-remote
- # ujoin group,eth0 — GRE: underlay mcast join for mGRE reception
- # forward/relay default to on, so they are omitted here.
- step(
- name,
- "nrlsmf debug 4 "
- f"instance smf-{name} "
- f"add overlay,cf,{GRE_DEV} "
- f"map {GRE_DEV},{cfg['underlay']},0.0.0.0 "
- f"ujoin {UNDERLAY_MCAST},eth0 "
- f"&> nrlsmf-mgre.log &",
- )
-
-for name in peer_names():
- wait_step(
- name,
- f'pgrep -af "nrlsmf.*instance smf-{name}"',
- match=f"smf-{name}",
- desc=f"{name} nrlsmf running on {GRE_DEV}",
- timeout=20,
- )
- wait_step(
- name,
- 'grep "regular group" nrlsmf-mgre.log',
- match="overlay",
- desc=f"{name} nrlsmf log shows overlay group",
- timeout=20,
- )
-
-section("Cleanup")
-
-for name in peer_names():
- step(name, "pkill nrlsmf || true")
- wait_step(
- name,
- "pgrep -af nrlsmf || true",
- match="",
- desc=f"{name} nrlsmf stopped",
- timeout=15,
- )
-
-test_step(True, "mGRE NBMA four-peer mutest completed")
diff --git a/tests/mutests/mgre_four_peers/mutest_mgre_mcast.py b/tests/mutests/mgre_four_peers/mutest_mgre_mcast.py
index 2ee3f22..01e117a 100644
--- a/tests/mutests/mgre_four_peers/mutest_mgre_mcast.py
+++ b/tests/mutests/mgre_four_peers/mutest_mgre_mcast.py
@@ -1,36 +1,80 @@
-"""Example: mGRE with multicast underlay remote + nrlsmf CF among four peers.
+"""Example: mGRE with multicast underlay remote + nrlsmf CF among four routers.
Topology (shared with other tests in this directory):
- p1 -- lan1 --\
- p2 -- lan2 ---\
- r0 (hub; underlay unicast + nrlsmf rmerge dataplane)
- p3 -- lan3 ---/
- p4 -- lan4 --/
+ h0 -- r0 -- lan0 --\\
+ h1 -- r1 -- lan1 ---\\
+ u0 (underlay: relays multicast between the LANs)
+ h2 -- r2 -- lan2 ---/
+ h3 -- r3 -- lan3 --/
+
+What "multicast-underlay mGRE" means here
+-------------------------------------------
+The other two mGRE modes in this directory (mutest_mgre_static.py and
+mutest_mgre_nhrp.py) both solve "which peer does this packet go to" by
+building a table -- static or dynamic -- that resolves each peer's
+overlay address to a specific underlay unicast address, then sending
+one unicast-encapsulated copy per peer. This mode does something
+different: instead of a table, the tunnel's *remote* address is
+configured as an IP multicast group address. A single encapsulated
+transmission is sent once, to that group, and the underlay network's
+own multicast routing fans it out to every router that has joined the
+group -- no per-peer table, no unicast replication at the sender.
+
+There's no real PIM multicast router available in this lab topology, so
+u0 stands in for "a multicast-capable underlay" by running nrlsmf
+itself in a pure dataplane role: `nrlsmf rmerge` across all four of its
+LAN-facing interfaces, which floods any multicast packet arriving on
+one of u0's interfaces out all the others. This is u0's *only* job in
+this test -- it does not participate in the GRE/mGRE overlay itself,
+and this nrlsmf instance on u0 is completely independent of the
+per-router overlay nrlsmf instances started later. A real deployment
+would use actual PIM multicast routing here instead of nrlsmf; this
+substitution exists purely to make the test self-contained without
+requiring a separate multicast routing daemon.
What this example covers
-------------------------
-* Underlay: unicast routing through r0, plus nrlsmf on r0 as a pure
- dataplane gateway (rmerge across eth0..eth3) so the GRE encapsulation
- group is flooded between the LANs. That instance is independent of the
- overlay nrlsmf instances on the peers.
-* Overlay: GRE tunnels with multicast remote (one underlay group reaches
- all peers). Peer nrlsmf classic flooding on mgre0 with map/ujoin.
-* Overlay multicast: iperf from p1 to p2/p3/p4 over the GRE overlay.
-
-See mutest_mgre.py for the NBMA (unicast underlay) multipoint GRE example.
+-------------------------
+* Underlay dataplane on u0: `nrlsmf rmerge` across eth0..eth3, standing
+ in for real underlay multicast routing (see above).
+* Overlay: GRE tunnels on r0..r3 with a multicast remote address
+ (mgre0), so one underlay multicast group reaches all four routers
+ symmetrically -- no hub, no per-peer replication.
+* Overlay nrlsmf: classic flooding (`cf`) on each router's host LAN
+ plus mgre0, with `ujoin` on each router's underlay eth0.
+* Overlay multicast from host h0, received at h1/h2/h3.
+
+See mutest_gre_p2p.py (point-to-point), mutest_mgre_static.py (static NBMA
+mGRE), mutest_mgre_nhrp.py (NHRP-resolved mGRE), and
+mutest_gre_external.py (external/metadata GRE) for the other GRE
+tunnel modes.
"""
+from munet.mutest.userapi import script_dir
from munet.mutest.userapi import section
from munet.mutest.userapi import step
from munet.mutest.userapi import test_step
from munet.mutest.userapi import wait_step
-PEERS = {
- "p1": {"underlay": "10.0.1.2", "overlay": "172.16.0.1"},
- "p2": {"underlay": "10.0.2.2", "overlay": "172.16.0.2"},
- "p3": {"underlay": "10.0.3.2", "overlay": "172.16.0.3"},
- "p4": {"underlay": "10.0.4.2", "overlay": "172.16.0.4"},
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from four_peer_hosts import RECV_HOSTS
+from four_peer_hosts import cleanup_iperf
+from four_peer_hosts import setup_host_lan
+from four_peer_hosts import start_host_mcast_client
+from four_peer_hosts import start_overlay_mcast_servers
+from four_peer_hosts import wait_overlay_mcast_receivers
+from smf_cli import check_common_show
+from smf_cli import check_show_neighbors
+from smf_cli import check_show_tunnel
+
+ROUTERS = {
+ "r0": {"underlay": "10.0.0.2", "overlay": "172.16.0.1"},
+ "r1": {"underlay": "10.0.1.2", "overlay": "172.16.0.2"},
+ "r2": {"underlay": "10.0.2.2", "overlay": "172.16.0.3"},
+ "r3": {"underlay": "10.0.3.2", "overlay": "172.16.0.4"},
}
UNDERLAY_MCAST = "239.1.1.1"
@@ -38,21 +82,17 @@
# Dedicated name: kernel fallback gre0 (remote any) steals the overlay
# subnet if addressed; see tunnel-setup comments below.
GRE_DEV = "mgre0"
-R0_IFACES = "eth0,eth1,eth2,eth3"
-
-
-def peer_names():
- return list(PEERS.keys())
+U0_IFACES = "eth0,eth1,eth2,eth3"
section("Disable offloads and wait for underlay addresses")
-step("r0", "sysctl -w net.ipv4.ip_forward=1")
-step("r0", "sysctl -w net.ipv4.conf.all.rp_filter=0")
-step("r0", "sysctl -w net.ipv4.conf.default.rp_filter=0")
-step("r0", "sysctl -w net.ipv4.conf.all.send_redirects=0")
+step("u0", "sysctl -w net.ipv4.ip_forward=1")
+step("u0", "sysctl -w net.ipv4.conf.all.rp_filter=0")
+step("u0", "sysctl -w net.ipv4.conf.default.rp_filter=0")
+step("u0", "sysctl -w net.ipv4.conf.all.send_redirects=0")
-for name, cfg in PEERS.items():
+for name, cfg in ROUTERS.items():
step(name, "ethtool -K eth0 rx off tx off || true")
step(name, "sysctl -w net.ipv4.conf.all.rp_filter=0")
step(name, "sysctl -w net.ipv4.conf.eth0.rp_filter=0")
@@ -66,25 +106,27 @@ def peer_names():
)
for ifname, addr in (
- ("eth0", "10.0.1.1"),
- ("eth1", "10.0.2.1"),
- ("eth2", "10.0.3.1"),
- ("eth3", "10.0.4.1"),
+ ("eth0", "10.0.0.1"),
+ ("eth1", "10.0.1.1"),
+ ("eth2", "10.0.2.1"),
+ ("eth3", "10.0.3.1"),
):
- step("r0", f"ethtool -K {ifname} rx off tx off || true")
- step("r0", f"sysctl -w net.ipv4.conf.{ifname}.rp_filter=0")
+ step("u0", f"ethtool -K {ifname} rx off tx off || true")
+ step("u0", f"sysctl -w net.ipv4.conf.{ifname}.rp_filter=0")
wait_step(
- "r0",
+ "u0",
f"ip -br addr show dev {ifname}",
match=addr,
- desc=f"r0 {ifname} address {addr}",
+ desc=f"u0 {ifname} address {addr}",
timeout=30,
)
-section("Underlay unicast reachability through hub router")
+setup_host_lan(step, wait_step)
-for src in peer_names():
- for dst, dcfg in PEERS.items():
+section("Underlay unicast reachability through u0")
+
+for src in ROUTERS:
+ for dst, dcfg in ROUTERS.items():
if src == dst:
continue
wait_step(
@@ -95,35 +137,37 @@ def peer_names():
timeout=20,
)
-section("Underlay mcast dataplane on r0 (nrlsmf rmerge)")
+section("Underlay multicast relay on u0 (stand-in for real PIM routing)")
-# Independent of peer overlay instances: flood multicast among the hub LANs
-# so GRE packets destined to UNDERLAY_MCAST reach every peer.
+# u0 is not part of the GRE/mGRE overlay -- this nrlsmf instance only
+# floods multicast between u0's four LAN interfaces, so GRE packets
+# addressed to UNDERLAY_MCAST reach every router. Stand-in for real
+# underlay multicast routing (e.g. PIM).
step(
- "r0",
+ "u0",
"nrlsmf debug 4 "
- "instance smf-r0-underlay "
- f"rmerge {R0_IFACES} "
- "&> nrlsmf-r0-underlay.log &",
+ "instance smf-u0-underlay "
+ f"rmerge {U0_IFACES} "
+ "&> nrlsmf-u0-underlay.log &",
)
wait_step(
- "r0",
- 'pgrep -af "nrlsmf.*instance smf-r0-underlay"',
- match="smf-r0-underlay",
- desc="r0 underlay nrlsmf running",
+ "u0",
+ 'pgrep -af "nrlsmf.*instance smf-u0-underlay"',
+ match="smf-u0-underlay",
+ desc="u0 underlay nrlsmf running",
timeout=20,
)
wait_step(
- "r0",
- 'grep "regular group" nrlsmf-r0-underlay.log',
+ "u0",
+ 'grep "regular group" nrlsmf-u0-underlay.log',
match="merge",
- desc="r0 nrlsmf log shows merge group",
+ desc="u0 nrlsmf log shows merge group",
timeout=20,
)
section("Create mGRE tunnels (multicast underlay remote)")
-for name, cfg in PEERS.items():
+for name, cfg in ROUTERS.items():
# Linux always creates a fallback gre0 (remote any / local any). If the
# overlay /24 lands on that device, routes prefer it over mgre0 and the
# multicast-remote tunnel never carries traffic. Flush and down gre0 so
@@ -133,15 +177,18 @@ def peer_names():
step(name, "ip addr flush dev gre0 2>/dev/null || true")
step(name, "ip link set gre0 down 2>/dev/null || true")
step(name, f"ip link del {GRE_DEV} 2>/dev/null || true")
- # Linux derives ikey/okey from the multicast remote; peers share that group.
+ # This is the defining difference from the other mGRE modes: remote
+ # is a multicast group address, not a unicast peer or 0.0.0.0.
+ # Linux derives ikey/okey from the multicast remote; peers share
+ # that group.
step(
name,
f"ip tunnel add {GRE_DEV} mode gre "
f"local {cfg['underlay']} remote {UNDERLAY_MCAST} ttl 64",
)
step(name, f"ip addr add {cfg['overlay']}/24 dev {GRE_DEV}")
+ step(name, f"ip link set {GRE_DEV} multicast on")
step(name, f"ip link set {GRE_DEV} up")
- step(name, f"ip route replace {OVERLAY_MCAST}/32 dev {GRE_DEV}")
wait_step(
name,
f"ip -br link show {GRE_DEV}",
@@ -155,10 +202,10 @@ def peer_names():
desc=f"{name} {GRE_DEV} remote is {UNDERLAY_MCAST}",
)
-section("Overlay unicast across mGRE (before peer nrlsmf)")
+section("Overlay unicast across mGRE (before overlay nrlsmf)")
-for src, scfg in PEERS.items():
- for dst, dcfg in PEERS.items():
+for src, scfg in ROUTERS.items():
+ for dst, dcfg in ROUTERS.items():
if src == dst:
continue
wait_step(
@@ -169,26 +216,18 @@ def peer_names():
timeout=20,
)
-section("Start peer nrlsmf classic flooding on mGRE (map + ujoin)")
+section("Start overlay nrlsmf classic flooding on mGRE (ujoin required)")
-for name, cfg in PEERS.items():
- # Overlay control/dataplane on peers (independent of r0 underlay instance):
- # instance smf-{name}-mcast — unique control pipe (/tmp shared across munet ns)
- # add overlay,cf,mgre0 — classic flooding on the GRE iface
- # map mgre0,local,0.0.0.0 — GRE: mGRE/any-remote mapping for SMF lookup
- # ujoin group,eth0 — GRE: join underlay mcast used for GRE encap/recv
- # forward/relay default to on, so they are omitted here.
+for name in ROUTERS:
step(
name,
"nrlsmf debug 4 "
f"instance smf-{name}-mcast "
- f"add overlay,cf,{GRE_DEV} "
- f"map {GRE_DEV},{cfg['underlay']},0.0.0.0 "
+ f"add overlay,cf,eth1,{GRE_DEV} "
+ f"layered {GRE_DEV} "
f"ujoin {UNDERLAY_MCAST},eth0 "
- f"&> nrlsmf-mgre-mcast.log &",
+ "&> nrlsmf-mgre-mcast.log &",
)
-
-for name in peer_names():
wait_step(
name,
f'pgrep -af "nrlsmf.*instance smf-{name}-mcast"',
@@ -204,44 +243,67 @@ def peer_names():
timeout=20,
)
-section("Overlay multicast through nrlsmf CF on mGRE")
-
-for name in ("p2", "p3", "p4"):
- step(
- name,
- f"iperf -u -T 4 -i 1 -s -e -B {OVERLAY_MCAST}%{GRE_DEV} "
- f"> iperf-mgre-server.log 2>&1 &",
+section("nrlsmf --cli show tunnel / neighbors (json, multicast-underlay remote)")
+
+check_common_show("u0", "smf-u0-underlay", group_name="merge")
+for name, cfg in ROUTERS.items():
+ inst = f"smf-{name}-mcast"
+ check_common_show(name, inst, group_name="overlay", ifaces=("eth1", GRE_DEV))
+ check_show_tunnel(
+ name, inst, GRE_DEV,
+ local=cfg["underlay"],
+ remotes=[UNDERLAY_MCAST],
+ overlay_ip=cfg["overlay"],
+ want_c=False,
)
+ # Device remote is the underlay group. NOARP, so no per-peer overlay neigh.
+ check_show_neighbors(
+ name, inst, GRE_DEV,
+ remotes=[UNDERLAY_MCAST],
+ min_count=1,
+ )
+
+section("[Mcast] Overlay multicast: h0 -> SMF -> h1/h2/h3")
step(
- "p1",
- f"iperf -u -T 4 -t 1000 -i 1 -b 8pps -l 1024 -e "
- f"-B {PEERS['p1']['overlay']}%{GRE_DEV} "
- f"-c {OVERLAY_MCAST} &> iperf-mgre-client.log &",
+ "r0",
+ f"tcpdump -l -n -i eth0 'proto gre or host {UNDERLAY_MCAST}' "
+ "> tcpdump-r0-eth0.log 2>&1 &",
)
-
-wait_step(
- "p1",
- "tail -n1 iperf-mgre-client.log",
- match="8 pps",
- desc="p1 sending overlay multicast at 8 pps",
- timeout=30,
+step(
+ "r0",
+ f"tcpdump -l -n -i {GRE_DEV} 'host {OVERLAY_MCAST}' "
+ "> tcpdump-r0-mgre0.log 2>&1 &",
+)
+step(
+ "r1",
+ f"tcpdump -l -n -i eth0 'proto gre or host {UNDERLAY_MCAST}' "
+ "> tcpdump-r1-eth0.log 2>&1 &",
)
+step(
+ "r1",
+ f"tcpdump -l -n -i {GRE_DEV} 'host {OVERLAY_MCAST}' "
+ "> tcpdump-r1-mgre0.log 2>&1 &",
+)
+step("r0", "sleep 1")
-for name in ("p2", "p3", "p4"):
- wait_step(
- name,
- "tail -n1 iperf-mgre-server.log",
- match="8 pps",
- desc=f"{name} receiving overlay multicast at 8 pps",
- timeout=45,
- )
+RECEIVERS = RECV_HOSTS
+start_overlay_mcast_servers(step, RECEIVERS, OVERLAY_MCAST)
+start_host_mcast_client(step, wait_step, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, RECEIVERS)
+
+step("r0", "pkill tcpdump || true")
+step("r1", "pkill tcpdump || true")
+step("r0", "echo '=== r0 eth0 ==='; cat tcpdump-r0-eth0.log || true")
+step("r0", "echo '=== r0 mgre0 ==='; cat tcpdump-r0-mgre0.log || true")
+step("r1", "echo '=== r1 eth0 ==='; cat tcpdump-r1-eth0.log || true")
+step("r1", "echo '=== r1 mgre0 ==='; cat tcpdump-r1-mgre0.log || true")
section("Cleanup")
-for name in peer_names():
+cleanup_iperf(step, RECEIVERS)
+for name in ROUTERS:
step(name, "pkill nrlsmf || true")
- step(name, "pkill iperf || true")
wait_step(
name,
"pgrep -af nrlsmf || true",
@@ -250,5 +312,5 @@ def peer_names():
timeout=15,
)
-step("r0", "pkill nrlsmf || true")
-test_step(True, "mGRE underlay-mcast four-peer mutest completed")
+step("u0", "pkill nrlsmf || true")
+test_step(True, "mGRE underlay-mcast four-router mutest completed")
diff --git a/tests/mutests/mgre_four_peers/mutest_mgre_nhrp.py b/tests/mutests/mgre_four_peers/mutest_mgre_nhrp.py
new file mode 100644
index 0000000..6244552
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/mutest_mgre_nhrp.py
@@ -0,0 +1,385 @@
+"""Example: NHRP-resolved mGRE (dynamic hub-and-spoke) + nrlsmf CF.
+
+Topology (shared with other tests in this directory):
+
+ h0 -- r0 -- lan0 --\\
+ r1 -- lan1 ---\\
+ u0 (underlay: routes ordinary IP between the LANs)
+ r2 -- lan2 ---/
+ r3 -- lan3 --/ <-- also the NHRP Server (NHS) / hub
+
+Why the NHS runs on r3, not on u0
+------------------------------------
+u0 represents the underlay/WAN -- the network that gets you from one
+router to another, but that the operator running the GRE/mGRE overlay
+generally does *not* control (think: the public Internet, a carrier
+MPLS core, or a satellite/cellular provider's network). It would be a
+mistake to put the NHRP Server there: the NHS is part of the overlay's
+own control plane, and needs to run somewhere the operator actually
+owns and trusts.
+
+So in this test, r3 -- one of the operator's own four routers, exactly
+like r0..r2 -- takes on the hub/NHS role in addition to being a normal
+overlay participant. r0..r2 are NHRP clients ("spokes") that register
+their real underlay address with r3 and are, from that point on,
+reachable by their overlay tunnel address without any manual mapping.
+u0 is not touched by NHRP configuration at all; it just routes plain
+IP packets between the four LANs, the same as in every other test in
+this directory.
+
+What "NHRP-resolved mGRE" means here
+---------------------------------------
+This is the third way of answering "which peer does this packet go to"
+on a multipoint GRE interface, and it's the one used by dynamic
+hub-and-spoke overlays such as Cisco-style DMVPN. Instead of a
+hand-built table (mutest_mgre_static.py), the overlay-to-underlay address
+table is built and kept up to date automatically by the Next Hop
+Resolution Protocol (NHRP, RFC 2332), via FRR's `nhrpd` daemon -- an
+external control-plane programming the tunnel. nrlsmf does not list
+peers; it uses `map …,dynamic` and reads the kernel `ip neigh` table
+nhrpd installed. Overlay unicast looks like static NBMA. Overlay
+multicast does not: a spoke only has the NHS in neigh, so the hub's
+nrlsmf is the replicator.
+
+What this example covers
+-------------------------
+* Underlay: unicast IP between all four routers, routed through u0 --
+ identical to the other tests. r3's own ordinary underlay address
+ (its `eth0` on lan3) is used directly as its NBMA identity; no
+ loopback or special addressing is needed on r3 or on u0.
+* Overlay: a single multipoint GRE interface (gre1) where overlay
+ unicast peer resolution is handled by `nhrpd`. Overlay tunnel
+ addresses (10.100.0.0/24) are a separate range from the underlay
+ addresses (10.0.0.0/24-10.0.3.0/24).
+* nrlsmf classic flooding (`cf`) on each router's host LAN plus gre1.
+ `map gre1,,dynamic` is the only inject config: nrlsmf learns
+ GRE dests from kernel `ip neigh` that nhrpd installed. That is the
+ point of this test -- nhrpd is an external tool programming the
+ tunnel, not nrlsmf listing peers itself.
+* Hub-and-spoke (DMVPN Phase 1), not spoke-to-spoke shortcuts. A spoke's
+ neigh table is the NHS; the hub's neigh table is every registered
+ spoke. Overlay multicast from h0 therefore goes r0 -> r3, and r3's
+ nrlsmf replicates to r1/r2. Spokes `layered gre1` so they do not
+ flood back out the tunnel; the hub is *not* layered, because it is
+ the overlay replicator. (Shortcuts need NFLOG/iptables and would not
+ fire on nrlsmf CF traffic anyway.)
+
+See mutest_gre_p2p.py (point-to-point), mutest_mgre_static.py (static NBMA
+mGRE), mutest_mgre_mcast.py (multicast-underlay mGRE), and
+mutest_gre_external.py (external/metadata GRE) for the other GRE
+tunnel modes.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from four_peer_hosts import RECV_HOSTS
+from four_peer_hosts import cleanup_iperf
+from four_peer_hosts import setup_host_lan
+from four_peer_hosts import start_host_mcast_client
+from four_peer_hosts import start_overlay_mcast_servers
+from four_peer_hosts import wait_overlay_mcast_receivers
+from smf_cli import check_common_show
+from smf_cli import check_show_neighbors
+from smf_cli import check_show_tunnel
+
+# Underlay LAN addressing (matches */etc.frr/frr.conf). Overlay/tunnel
+# addresses are deliberately drawn from a separate 10.100.0.0/24 range
+# so they can never collide with underlay addresses. r3 is the hub/NHS;
+# r0..r2 are spokes.
+HUB = "r3"
+SPOKES = ("r0", "r1", "r2")
+
+ROUTERS = {
+ "r0": {"underlay": "10.0.0.2", "overlay": "10.100.0.2"},
+ "r1": {"underlay": "10.0.1.2", "overlay": "10.100.0.3"},
+ "r2": {"underlay": "10.0.2.2", "overlay": "10.100.0.4"},
+ "r3": {"underlay": "10.0.3.2", "overlay": "10.100.0.1"}, # hub/NHS
+}
+
+GRE_DEV = "gre1"
+NHRP_NETWORK_ID = "1"
+GRE_KEY = "42"
+OVERLAY_MCAST = "239.0.0.1"
+
+
+section("Disable offloads and wait for underlay addresses")
+
+for name, cfg in ROUTERS.items():
+ step(name, "ethtool -K eth0 rx off tx off || true")
+ wait_step(
+ name,
+ "ip -br addr show dev eth0",
+ match=cfg["underlay"],
+ desc=f"{name} underlay address {cfg['underlay']}",
+ timeout=30,
+ )
+
+for ifname, addr in (
+ ("eth0", "10.0.0.1"),
+ ("eth1", "10.0.1.1"),
+ ("eth2", "10.0.2.1"),
+ ("eth3", "10.0.3.1"),
+):
+ step("u0", f"ethtool -K {ifname} rx off tx off || true")
+ wait_step(
+ "u0",
+ f"ip -br addr show dev {ifname}",
+ match=addr,
+ desc=f"u0 {ifname} address {addr}",
+ timeout=30,
+ )
+
+setup_host_lan(step, wait_step)
+
+step("u0", "sysctl -w net.ipv4.ip_forward=1")
+
+section("Underlay unicast reachability (spokes to hub, through u0)")
+
+# u0 has no NHRP role and no special configuration here -- it's just
+# plain IP transit, exactly as in every other test in this directory.
+for name in SPOKES:
+ wait_step(
+ name,
+ f"ping -c1 -W2 {ROUTERS[HUB]['underlay']}",
+ match="1 received",
+ desc=f"{name} ping hub underlay {HUB} ({ROUTERS[HUB]['underlay']})",
+ timeout=20,
+ )
+
+section("Create multipoint GRE interface on the hub and each spoke")
+
+# Kernel fallback gre0 (remote any) will steal inbound GRE if left up.
+step(HUB, "ip addr flush dev gre0 2>/dev/null || true")
+step(HUB, "ip link set gre0 down 2>/dev/null || true")
+step(HUB, f"ip link del {GRE_DEV} 2>/dev/null || true")
+step(
+ HUB,
+ f"ip tunnel add {GRE_DEV} mode gre "
+ f"local {ROUTERS[HUB]['underlay']} key {GRE_KEY} ttl 64",
+)
+step(HUB, f"ip addr add {ROUTERS[HUB]['overlay']}/32 dev {GRE_DEV}")
+step(HUB, f"ip link set {GRE_DEV} multicast on")
+step(HUB, f"ip link set {GRE_DEV} up")
+wait_step(
+ HUB,
+ f"ip -br link show {GRE_DEV}",
+ match="UP",
+ desc=f"{HUB} {GRE_DEV} is UP",
+)
+
+for name in SPOKES:
+ cfg = ROUTERS[name]
+ step(name, "ip addr flush dev gre0 2>/dev/null || true")
+ step(name, "ip link set gre0 down 2>/dev/null || true")
+ step(name, f"ip link del {GRE_DEV} 2>/dev/null || true")
+ step(
+ name,
+ f"ip tunnel add {GRE_DEV} mode gre "
+ f"local {cfg['underlay']} key {GRE_KEY} ttl 64",
+ )
+ step(name, f"ip addr add {cfg['overlay']}/32 dev {GRE_DEV}")
+ step(name, f"ip link set {GRE_DEV} multicast on")
+ step(name, f"ip link set {GRE_DEV} up")
+ wait_step(
+ name,
+ f"ip -br link show {GRE_DEV}",
+ match="UP",
+ desc=f"{name} {GRE_DEV} is UP",
+ )
+
+section("Configure FRR nhrpd: hub (NHS) on r3")
+
+wait_step(
+ HUB,
+ "pgrep -af nhrpd",
+ match="nhrpd",
+ desc=f"{HUB} nhrpd is running",
+ timeout=20,
+)
+
+step(
+ HUB,
+ f"vtysh -c 'configure terminal' "
+ f"-c 'interface {GRE_DEV}' "
+ f"-c 'ip nhrp network-id {NHRP_NETWORK_ID}' "
+ f"-c 'ip nhrp registration no-unique'",
+)
+
+section("Configure FRR nhrpd: spokes (NHC, register with r3)")
+
+for name in SPOKES:
+ wait_step(
+ name,
+ "pgrep -af nhrpd",
+ match="nhrpd",
+ desc=f"{name} nhrpd is running",
+ timeout=20,
+ )
+ step(
+ name,
+ f"vtysh -c 'configure terminal' "
+ f"-c 'interface {GRE_DEV}' "
+ f"-c 'ip nhrp network-id {NHRP_NETWORK_ID}' "
+ f"-c 'ip nhrp nhs {ROUTERS[HUB]['overlay']} "
+ f"nbma {ROUTERS[HUB]['underlay']}' "
+ f"-c 'ip nhrp registration no-unique'",
+ )
+
+section("Wait for spoke registration with the hub")
+
+# Registration is sent shortly after the nhs is configured; give it
+# room to converge rather than asserting on internal timer behavior.
+for name in SPOKES:
+ cfg = ROUTERS[name]
+ wait_step(
+ HUB,
+ "vtysh -c 'show ip nhrp cache'",
+ match=cfg["overlay"],
+ desc=f"{HUB} NHRP cache shows {name} registered ({cfg['overlay']})",
+ timeout=60,
+ )
+
+for name in SPOKES:
+ wait_step(
+ name,
+ "vtysh -c 'show ip nhrp nhs'",
+ match=ROUTERS[HUB]["overlay"],
+ desc=f"{name} NHRP shows hub NHS {ROUTERS[HUB]['overlay']}",
+ timeout=30,
+ )
+
+section("Overlay unicast, hub <-> spokes, resolved via NHRP (before nrlsmf)")
+
+for name in SPOKES:
+ cfg = ROUTERS[name]
+ wait_step(
+ name,
+ f"ping -c1 -W3 {ROUTERS[HUB]['overlay']}",
+ match="1 received",
+ desc=f"{name} ping hub overlay {ROUTERS[HUB]['overlay']}",
+ timeout=20,
+ )
+ wait_step(
+ HUB,
+ f"ping -c1 -W3 {cfg['overlay']}",
+ match="1 received",
+ desc=f"{HUB} ping {name} overlay ({cfg['overlay']})",
+ timeout=20,
+ )
+
+section("Kernel neigh that nhrpd programmed (nrlsmf map dynamic source)")
+
+# Spoke: NHS only. Hub: every registered spoke. Overlay pings above
+# make sure those entries are in the kernel, not only in nhrpd's cache.
+for name in SPOKES:
+ wait_step(
+ name,
+ f"ip neigh show dev {GRE_DEV}",
+ match=ROUTERS[HUB]["overlay"],
+ desc=f"{name} gre1 neigh has hub {ROUTERS[HUB]['overlay']}",
+ timeout=20,
+ )
+for name in SPOKES:
+ wait_step(
+ HUB,
+ f"ip neigh show dev {GRE_DEV}",
+ match=ROUTERS[name]["overlay"],
+ desc=f"{HUB} gre1 neigh has {name} {ROUTERS[name]['overlay']}",
+ timeout=20,
+ )
+
+section("Start nrlsmf (map dynamic from nhrpd neigh; hub replicates)")
+
+# Spokes are layered so a packet received on gre1 is not flooded back
+# out of it. The hub is not: it is the overlay-mcast replicator, using
+# the spoke dests nhrpd installed.
+for name, cfg in ROUTERS.items():
+ layered = f"layered {GRE_DEV} " if name in SPOKES else ""
+ step(
+ name,
+ "nrlsmf debug 4 "
+ f"instance smf-{name} "
+ f"add overlay,cf,eth1,{GRE_DEV} "
+ f"{layered}"
+ f"map {GRE_DEV},{cfg['underlay']},dynamic "
+ "&> nrlsmf-mgre-nhrp.log &",
+ )
+ wait_step(
+ name,
+ f'pgrep -af "nrlsmf.*instance smf-{name}"',
+ match=f"smf-{name}",
+ desc=f"{name} nrlsmf running on {GRE_DEV}",
+ timeout=20,
+ )
+ wait_step(
+ name,
+ 'grep "regular group" nrlsmf-mgre-nhrp.log',
+ match="overlay",
+ desc=f"{name} nrlsmf log shows overlay group",
+ timeout=20,
+ )
+
+section("nrlsmf --cli show tunnel / neighbors (json, NHRP map dynamic)")
+
+# Spokes see the NHS; the hub sees every registered spoke. Overlay Neighbor IP
+# comes from nhrpd's kernel neigh; C is not required (dynamic, not explicit map).
+for name in SPOKES:
+ inst = f"smf-{name}"
+ check_common_show(name, inst, group_name="overlay", ifaces=("eth1", GRE_DEV))
+ check_show_tunnel(
+ name, inst, GRE_DEV,
+ local=ROUTERS[name]["underlay"],
+ remotes=[ROUTERS[HUB]["underlay"]],
+ overlay_ip=ROUTERS[name]["overlay"],
+ )
+ check_show_neighbors(
+ name, inst, GRE_DEV,
+ remotes=[ROUTERS[HUB]["underlay"]],
+ neighbor_ips=[ROUTERS[HUB]["overlay"]],
+ min_count=1,
+ )
+
+hub_inst = f"smf-{HUB}"
+check_common_show(HUB, hub_inst, group_name="overlay", ifaces=("eth1", GRE_DEV))
+check_show_tunnel(
+ HUB, hub_inst, GRE_DEV,
+ local=ROUTERS[HUB]["underlay"],
+ remotes=[ROUTERS[s]["underlay"] for s in SPOKES],
+ overlay_ip=ROUTERS[HUB]["overlay"],
+)
+check_show_neighbors(
+ HUB, hub_inst, GRE_DEV,
+ remotes=[ROUTERS[s]["underlay"] for s in SPOKES],
+ neighbor_ips=[ROUTERS[s]["overlay"] for s in SPOKES],
+ min_count=3,
+)
+
+section("[NHRP] Overlay multicast: h0 -> r0 -> hub -> h1/h2/h3")
+
+RECEIVERS = RECV_HOSTS
+start_overlay_mcast_servers(step, RECEIVERS, OVERLAY_MCAST)
+start_host_mcast_client(step, wait_step, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, RECEIVERS)
+
+section("Cleanup")
+
+cleanup_iperf(step, RECEIVERS)
+for name in ROUTERS:
+ step(name, "pkill nrlsmf || true")
+ wait_step(
+ name,
+ "pgrep -af nrlsmf || true",
+ match="",
+ desc=f"{name} nrlsmf stopped",
+ timeout=15,
+ )
+
+test_step(True, "NHRP-resolved mGRE four-router mutest completed")
diff --git a/tests/mutests/mgre_four_peers/mutest_mgre_static.py b/tests/mutests/mgre_four_peers/mutest_mgre_static.py
new file mode 100644
index 0000000..845b47f
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/mutest_mgre_static.py
@@ -0,0 +1,393 @@
+"""Example: multipoint GRE (static NBMA) + nrlsmf CF among four routers.
+
+Topology (shared with other tests in this directory):
+
+ +-------+
+ h0 -- r0 -- lan0 -----+ |
+ h1 -- r1 -- lan1 -----+ u0 | (underlay: routes ordinary IP between the LANs)
+ h2 -- r2 -- lan2 -----+ |
+ h3 -- r3 -- lan3 -----+ |
+ +-------+
+
+What "static NBMA mGRE" means here
+------------------------------------
+A single mGRE interface (kernel `remote 0.0.0.0`) can represent many
+remote peers instead of just one, but GRE itself has no way to say
+"which peer does this packet go to" -- something else has to answer
+that. Here, that something is a plain, hand-built table: each router's
+kernel neighbor table (`ip neigh`) maps every other router's *overlay*
+tunnel address to that router's *underlay* address, the same way an
+ARP table maps an IP to a MAC address on an Ethernet LAN. It's manual
+and doesn't scale to large or changing meshes, but it needs no extra
+daemon or protocol -- just entries you set once.
+
+(See mutest_mgre_nhrp.py for the alternative where this same table is
+filled in dynamically by a routing daemon instead of by hand -- at the
+kernel level the two are the same mechanism, just maintained
+differently.)
+
+What this example covers
+-------------------------
+* Underlay: unicast IP between all four routers, routed through u0.
+* Overlay: one shared multipoint GRE interface (gre1) on each router,
+ all in the same 172.16.0.0/24 overlay subnet.
+* Two peer tables, because overlay unicast and overlay multicast are
+ resolved differently:
+ - Kernel `ip neigh`: overlay unicast address -> underlay address.
+ Used by the GRE driver for ordinary overlay ping/unicast.
+ - nrlsmf overlay-multicast inject dests, two passes:
+ 1. Explicit `map gre1,,` for every other router.
+ 2. Stop nrlsmf on sender r0 and restart it with
+ `map gre1,,dynamic` so the same dests are learned
+ from kernel `ip neigh`. Receivers keep their explicit maps.
+ `0.0.0.0` remains the kernel wildcard remote (not a send dest).
+* nrlsmf classic flooding (`cf`) on each router's host LAN plus gre1.
+ Iperf is sourced on h0 and received on h1/h2/h3.
+
+After the cf overlay-mcast passes, the same nrlsmf processes get
+runtime ``with-frr`` and ``elastic overlay`` (maps and grouping stay).
+Keep h1 receiving, stop h2/h3, and expect at most 1 pps on the
+stopped hosts.
+
+See mutest_gre_p2p.py (point-to-point), mutest_mgre_nhrp.py (NHRP-
+resolved mGRE), mutest_mgre_mcast.py (multicast-underlay mGRE), and
+mutest_gre_external.py (external/metadata GRE, resolved via routes
+instead of this file's neighbor table) for the other GRE tunnel modes.
+"""
+
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+import time
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from four_peer_hosts import RECV_HOSTS
+from four_peer_hosts import ROUTER_HOST_IFACE
+from four_peer_hosts import SOURCE_HOST
+from four_peer_hosts import cleanup_iperf
+from four_peer_hosts import count_overlay_mcast_pkts
+from four_peer_hosts import enable_host_igmp
+from four_peer_hosts import restart_overlay_mcast_servers
+from four_peer_hosts import setup_host_lan
+from four_peer_hosts import start_host_mcast_client
+from four_peer_hosts import start_overlay_mcast_servers
+from four_peer_hosts import wait_igmp_group
+from four_peer_hosts import wait_overlay_mcast_receivers
+from smf_cli import check_common_show
+from smf_cli import check_group_elastic
+from smf_cli import check_show_neighbors
+from smf_cli import check_show_tunnel
+from smf_cli import cli_cmd
+from smf_cli import wait_em_flow
+from smf_cli import wait_em_managed
+
+# Underlay LAN addressing (matches */etc.frr/frr.conf)
+ROUTERS = {
+ "r0": {"underlay": "10.0.0.2", "overlay": "172.16.0.1"},
+ "r1": {"underlay": "10.0.1.2", "overlay": "172.16.0.2"},
+ "r2": {"underlay": "10.0.2.2", "overlay": "172.16.0.3"},
+ "r3": {"underlay": "10.0.3.2", "overlay": "172.16.0.4"},
+}
+
+OVERLAY_MCAST = "239.0.0.1"
+# Dedicated name: kernel fallback gre0 (remote any) will steal inbound
+# GRE and the overlay subnet if we try to reuse it.
+GRE_DEV = "gre1"
+
+
+section("Disable offloads and wait for underlay addresses")
+
+step("u0", "sysctl -w net.ipv4.ip_forward=1")
+
+for name, cfg in ROUTERS.items():
+ step(name, "ethtool -K eth0 rx off tx off || true")
+ wait_step(
+ name,
+ "ip -br addr show dev eth0",
+ match=cfg["underlay"],
+ desc=f"{name} underlay address {cfg['underlay']}",
+ timeout=30,
+ )
+
+for ifname, addr in (
+ ("eth0", "10.0.0.1"),
+ ("eth1", "10.0.1.1"),
+ ("eth2", "10.0.2.1"),
+ ("eth3", "10.0.3.1"),
+):
+ step("u0", f"ethtool -K {ifname} rx off tx off || true")
+ wait_step(
+ "u0",
+ f"ip -br addr show dev {ifname}",
+ match=addr,
+ desc=f"u0 {ifname} address {addr}",
+ timeout=30,
+ )
+
+setup_host_lan(step, wait_step)
+
+section("Underlay unicast reachability through u0")
+
+for src in ROUTERS:
+ for dst, dcfg in ROUTERS.items():
+ if src == dst:
+ continue
+ wait_step(
+ src,
+ f"ping -c1 -W2 {dcfg['underlay']}",
+ match="1 received",
+ desc=f"{src} ping underlay {dst} ({dcfg['underlay']})",
+ timeout=20,
+ )
+
+section("Create multipoint GRE tunnels with static NBMA neighbors")
+
+for name, cfg in ROUTERS.items():
+ step(name, "ip addr flush dev gre0 2>/dev/null || true")
+ step(name, "ip link set gre0 down 2>/dev/null || true")
+ step(name, f"ip link del {GRE_DEV} 2>/dev/null || true")
+ # remote 0.0.0.0: this is what makes the interface multipoint
+ # rather than point-to-point -- there's no single fixed peer.
+ step(
+ name,
+ f"ip link add name {GRE_DEV} type gre "
+ f"local {cfg['underlay']} remote 0.0.0.0 ttl 64",
+ )
+ step(name, f"ip addr add {cfg['overlay']}/24 dev {GRE_DEV}")
+ step(name, f"ip link set {GRE_DEV} multicast on")
+ step(name, f"ip link set {GRE_DEV} up")
+ # Static NBMA table: map every other router's overlay address to
+ # its underlay address, by hand. This is the "static" half of
+ # "static NBMA mGRE" -- an NHRP daemon would maintain these same
+ # kinds of entries dynamically instead (see mutest_mgre_nhrp.py).
+ for other, ocfg in ROUTERS.items():
+ if other == name:
+ continue
+ step(
+ name,
+ f"ip neigh replace {ocfg['overlay']} lladdr {ocfg['underlay']} "
+ f"nud permanent dev {GRE_DEV}",
+ )
+ wait_step(
+ name,
+ f"ip -br link show {GRE_DEV}",
+ match="UP",
+ desc=f"{name} {GRE_DEV} is UP",
+ )
+
+section("Overlay unicast across mGRE (before nrlsmf)")
+
+for src, scfg in ROUTERS.items():
+ for dst, dcfg in ROUTERS.items():
+ if src == dst:
+ continue
+ wait_step(
+ src,
+ f"ping -c1 -W3 -I {scfg['overlay']} {dcfg['overlay']}",
+ match="1 received",
+ desc=f"{src} ping overlay {dst} ({dcfg['overlay']})",
+ timeout=20,
+ )
+
+section("Start nrlsmf classic flooding on mGRE")
+
+# Overlay mcast inject dests: one map per other router's underlay address.
+for name, cfg in ROUTERS.items():
+ maps = " ".join(
+ f"map {GRE_DEV},{cfg['underlay']},{ocfg['underlay']}"
+ for other, ocfg in ROUTERS.items()
+ if other != name
+ )
+ step(
+ name,
+ "nrlsmf debug 4 "
+ f"instance smf-{name} "
+ f"add overlay,cf,eth1,{GRE_DEV} "
+ f"layered {GRE_DEV} "
+ f"{maps} "
+ "&> nrlsmf-mgre.log &",
+ )
+ wait_step(
+ name,
+ f'pgrep -af "nrlsmf.*instance smf-{name}"',
+ match=f"smf-{name}",
+ desc=f"{name} nrlsmf running on {GRE_DEV}",
+ timeout=20,
+ )
+ wait_step(
+ name,
+ 'grep "regular group" nrlsmf-mgre.log',
+ match="overlay",
+ desc=f"{name} nrlsmf log shows overlay group",
+ timeout=20,
+ )
+
+section("nrlsmf --cli show tunnel / neighbors (json, explicit map)")
+
+for name, cfg in ROUTERS.items():
+ peers = [ocfg for other, ocfg in ROUTERS.items() if other != name]
+ inst = f"smf-{name}"
+ check_common_show(name, inst, group_name="overlay", ifaces=("eth1", GRE_DEV))
+ check_show_tunnel(
+ name, inst, GRE_DEV,
+ local=cfg["underlay"],
+ remotes=[p["underlay"] for p in peers],
+ overlay_ip=cfg["overlay"],
+ want_c=True,
+ )
+ check_show_neighbors(
+ name, inst, GRE_DEV,
+ remotes=[p["underlay"] for p in peers],
+ neighbor_ips=[p["overlay"] for p in peers],
+ min_count=3,
+ want_c=True,
+ )
+
+section("[Static] Overlay multicast: h0 -> SMF -> h1/h2/h3 (explicit map)")
+
+RECEIVERS = RECV_HOSTS
+start_overlay_mcast_servers(step, RECEIVERS, OVERLAY_MCAST)
+start_host_mcast_client(step, wait_step, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, RECEIVERS)
+
+section("[Static] Overlay multicast: h0 -> SMF -> h1/h2/h3 (r0 map dynamic)")
+
+cleanup_iperf(step, RECEIVERS)
+step("r0", "pkill nrlsmf || true")
+wait_step(
+ "r0",
+ "pgrep -af nrlsmf || true",
+ match="",
+ desc="r0 nrlsmf stopped",
+ timeout=15,
+)
+step(
+ "r0",
+ "nrlsmf debug 4 "
+ "instance smf-r0 "
+ f"add overlay,cf,eth1,{GRE_DEV} "
+ f"layered {GRE_DEV} "
+ f"map {GRE_DEV},{ROUTERS['r0']['underlay']},dynamic "
+ "&> nrlsmf-mgre-dynamic.log &",
+)
+wait_step(
+ "r0",
+ 'pgrep -af "nrlsmf.*instance smf-r0"',
+ match="smf-r0",
+ desc="r0 nrlsmf running on gre1 (map dynamic)",
+ timeout=20,
+)
+wait_step(
+ "r0",
+ 'grep "regular group" nrlsmf-mgre-dynamic.log',
+ match="overlay",
+ desc="r0 nrlsmf log shows overlay group",
+ timeout=20,
+)
+
+section("nrlsmf --cli show tunnel / neighbors (json, r0 map dynamic)")
+
+# Dynamic learns remotes from kernel neigh; Neighbor IP is overlay, Remote is
+# underlay. C is only set for explicit map, so it is not required here.
+peers_r0 = [ocfg for other, ocfg in ROUTERS.items() if other != "r0"]
+check_show_tunnel(
+ "r0", "smf-r0", GRE_DEV,
+ local=ROUTERS["r0"]["underlay"],
+ remotes=[p["underlay"] for p in peers_r0],
+ overlay_ip=ROUTERS["r0"]["overlay"],
+)
+check_show_neighbors(
+ "r0", "smf-r0", GRE_DEV,
+ remotes=[p["underlay"] for p in peers_r0],
+ neighbor_ips=[p["overlay"] for p in peers_r0],
+ min_count=3,
+)
+
+start_overlay_mcast_servers(step, RECEIVERS, OVERLAY_MCAST)
+start_host_mcast_client(step, wait_step, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, RECEIVERS)
+
+KEEP_RECV = ("h1",)
+IDLE_RECV = ("h2", "h3")
+EM_WINDOW_S = 2.0
+EM_IDLE_MAX_PPS = 1
+
+
+section("[Static] Host IGMP on eth1; h0 joins 239.0.0.1")
+
+# h0 is the sender; join 239.0.0.1 the same way h1/h2/h3 do so r0
+# eth1 is an IGMP host LAN and with-frr sees a local member.
+start_overlay_mcast_servers(step, (SOURCE_HOST,), OVERLAY_MCAST)
+enable_host_igmp(step, wait_step, ROUTERS)
+for name in ROUTERS:
+ wait_igmp_group(wait_step, name, OVERLAY_MCAST)
+
+section("[Static] Runtime with-frr + elastic overlay")
+
+# Same running CF processes. Maps, layered gre1, and overlay grouping
+# stay. with-frr starts FRR IGMP polling; elastic overlay turns on EM.
+for name in ROUTERS:
+ inst = f"smf-{name}"
+ cli_cmd(name, "with-frr", inst)
+ cli_cmd(name, "elastic overlay", inst)
+ check_group_elastic(
+ name, inst, "overlay", enabled=True,
+ ifaces=(ROUTER_HOST_IFACE, GRE_DEV),
+ )
+
+section("[Static] EM ack/forwarding state")
+
+# r1 has a local IGMP member (h1): it must ACK upstream so r0 FORWARDs
+# on gre1. Default FwdStatus is usually LIMIT; oif is the per-iface
+# token-bucket status after EM_ACK (show groups details).
+wait_em_flow("r1", "smf-r1", OVERLAY_MCAST, ack="yes", status="ACTIVE", timeout=20)
+wait_em_flow("r0", "smf-r0", OVERLAY_MCAST, ack="yes", timeout=20)
+wait_em_managed("r1", "smf-r1", ROUTER_HOST_IFACE, OVERLAY_MCAST, timeout=20)
+
+section("[Static] Overlay mcast after elastic")
+# New h1 iperf log so grep cannot match CF-era lines.
+restart_overlay_mcast_servers(step, KEEP_RECV, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, KEEP_RECV)
+
+section("[Static] Stop h2/h3; keep h1")
+
+for name in IDLE_RECV:
+ step(name, "pkill iperf || true")
+
+time.sleep(6)
+wait_step(
+ "h1",
+ "tail -n1 iperf-h1-server.log",
+ match="8 pps",
+ desc="h1 still receiving application multicast at 8 pps after elastic",
+ timeout=30,
+)
+
+for name in IDLE_RECV:
+ n = count_overlay_mcast_pkts(step, name, OVERLAY_MCAST, window_s=EM_WINDOW_S)
+ pps = n / EM_WINDOW_S
+ test_step(
+ pps <= EM_IDLE_MAX_PPS,
+ f"{name} non-receiver overlay mcast {pps:.1f} pps (at most {EM_IDLE_MAX_PPS})",
+ target=name,
+ )
+
+section("Cleanup")
+
+cleanup_iperf(step, RECEIVERS)
+for name in ROUTERS:
+ step(name, "pkill nrlsmf || true")
+ wait_step(
+ name,
+ "pgrep -af nrlsmf || true",
+ match="",
+ desc=f"{name} nrlsmf stopped",
+ timeout=15,
+ )
+
+test_step(True, "mGRE NBMA four-router mutest completed")
diff --git a/tests/mutests/mgre_four_peers/p1/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/p1/etc.frr/frr.conf
deleted file mode 100644
index 3283210..0000000
--- a/tests/mutests/mgre_four_peers/p1/etc.frr/frr.conf
+++ /dev/null
@@ -1,7 +0,0 @@
-log file /var/log/frr/frr.log
-!
-ip route 0.0.0.0/0 10.0.1.1
-!
-interface eth0
- ip address 10.0.1.2/24
-!
diff --git a/tests/mutests/mgre_four_peers/p2/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/p2/etc.frr/frr.conf
deleted file mode 100644
index 3602427..0000000
--- a/tests/mutests/mgre_four_peers/p2/etc.frr/frr.conf
+++ /dev/null
@@ -1,7 +0,0 @@
-log file /var/log/frr/frr.log
-!
-ip route 0.0.0.0/0 10.0.2.1
-!
-interface eth0
- ip address 10.0.2.2/24
-!
diff --git a/tests/mutests/mgre_four_peers/p3/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/p3/etc.frr/frr.conf
deleted file mode 100644
index a188bad..0000000
--- a/tests/mutests/mgre_four_peers/p3/etc.frr/frr.conf
+++ /dev/null
@@ -1,7 +0,0 @@
-log file /var/log/frr/frr.log
-!
-ip route 0.0.0.0/0 10.0.3.1
-!
-interface eth0
- ip address 10.0.3.2/24
-!
diff --git a/tests/mutests/mgre_four_peers/p4/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/p4/etc.frr/frr.conf
deleted file mode 100644
index 2bf878a..0000000
--- a/tests/mutests/mgre_four_peers/p4/etc.frr/frr.conf
+++ /dev/null
@@ -1,7 +0,0 @@
-log file /var/log/frr/frr.log
-!
-ip route 0.0.0.0/0 10.0.4.1
-!
-interface eth0
- ip address 10.0.4.2/24
-!
diff --git a/tests/mutests/mgre_four_peers/r0/etc.frr/daemons b/tests/mutests/mgre_four_peers/r0/etc.frr/daemons
index e69de29..6d87a20 100644
--- a/tests/mutests/mgre_four_peers/r0/etc.frr/daemons
+++ b/tests/mutests/mgre_four_peers/r0/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_four_peers/r0/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/r0/etc.frr/frr.conf
index c29acbd..dc81c13 100644
--- a/tests/mutests/mgre_four_peers/r0/etc.frr/frr.conf
+++ b/tests/mutests/mgre_four_peers/r0/etc.frr/frr.conf
@@ -1,17 +1,14 @@
log file /var/log/frr/frr.log
!
-! Hub router: four underlay LANs with static (connected) routes.
-! Underlay multicast for mGRE is forwarded by smcrouted (see mutest).
+! r0: SMF router. eth0 is the underlay toward u0; eth1 is the host
+! LAN toward h0 (iperf source).
+!
+ip route 0.0.0.0/0 10.0.0.1
!
interface eth0
- ip address 10.0.1.1/24
+ ip address 10.0.0.2/24
!
interface eth1
- ip address 10.0.2.1/24
-!
-interface eth2
- ip address 10.0.3.1/24
-!
-interface eth3
- ip address 10.0.4.1/24
+ ip address 192.168.55.1/24
+ ip igmp
!
diff --git a/tests/mutests/mgre_four_peers/r1/etc.frr/daemons b/tests/mutests/mgre_four_peers/r1/etc.frr/daemons
new file mode 100644
index 0000000..6d87a20
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r1/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_four_peers/r1/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/r1/etc.frr/frr.conf
new file mode 100644
index 0000000..91b6cb4
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r1/etc.frr/frr.conf
@@ -0,0 +1,14 @@
+log file /var/log/frr/frr.log
+!
+! r1: SMF router. eth0 is the underlay toward u0; eth1 is the host
+! LAN toward h1 (iperf receiver).
+!
+ip route 0.0.0.0/0 10.0.1.1
+!
+interface eth0
+ ip address 10.0.1.2/24
+!
+interface eth1
+ ip address 192.168.56.1/24
+ ip igmp
+!
diff --git a/tests/mutests/mgre_four_peers/r1/etc.frr/vtysh.conf b/tests/mutests/mgre_four_peers/r1/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r1/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_four_peers/r2/etc.frr/daemons b/tests/mutests/mgre_four_peers/r2/etc.frr/daemons
new file mode 100644
index 0000000..6d87a20
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r2/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_four_peers/r2/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/r2/etc.frr/frr.conf
new file mode 100644
index 0000000..bd6bd6e
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r2/etc.frr/frr.conf
@@ -0,0 +1,14 @@
+log file /var/log/frr/frr.log
+!
+! r2: SMF router. eth0 is the underlay toward u0; eth1 is the host
+! LAN toward h2 (iperf receiver).
+!
+ip route 0.0.0.0/0 10.0.2.1
+!
+interface eth0
+ ip address 10.0.2.2/24
+!
+interface eth1
+ ip address 192.168.57.1/24
+ ip igmp
+!
diff --git a/tests/mutests/mgre_four_peers/r2/etc.frr/vtysh.conf b/tests/mutests/mgre_four_peers/r2/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r2/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_four_peers/r3/etc.frr/daemons b/tests/mutests/mgre_four_peers/r3/etc.frr/daemons
new file mode 100644
index 0000000..6d87a20
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r3/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_four_peers/r3/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/r3/etc.frr/frr.conf
new file mode 100644
index 0000000..df4b5c4
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r3/etc.frr/frr.conf
@@ -0,0 +1,16 @@
+log file /var/log/frr/frr.log
+!
+! r3: SMF router. eth0 is the underlay toward u0; eth1 is the host
+! LAN toward h3 (iperf receiver). In the NHRP mutest, r3 additionally
+! takes on the hub/NHRP-Server (NHS) role -- always on an
+! operator-controlled router, never on the untrusted underlay (u0).
+!
+ip route 0.0.0.0/0 10.0.3.1
+!
+interface eth0
+ ip address 10.0.3.2/24
+!
+interface eth1
+ ip address 192.168.58.1/24
+ ip igmp
+!
diff --git a/tests/mutests/mgre_four_peers/r3/etc.frr/vtysh.conf b/tests/mutests/mgre_four_peers/r3/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/r3/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_four_peers/u0/etc.frr/daemons b/tests/mutests/mgre_four_peers/u0/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_four_peers/u0/etc.frr/frr.conf b/tests/mutests/mgre_four_peers/u0/etc.frr/frr.conf
new file mode 100644
index 0000000..03da9ed
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/u0/etc.frr/frr.conf
@@ -0,0 +1,18 @@
+log file /var/log/frr/frr.log
+!
+! Underlay ("u0"): routes ordinary IP traffic between the four LANs.
+! Plays no active GRE/nrlsmf role except in the multicast-underlay test,
+! where it relays multicast between the LANs as a pure dataplane helper.
+!
+interface eth0
+ ip address 10.0.0.1/24
+!
+interface eth1
+ ip address 10.0.1.1/24
+!
+interface eth2
+ ip address 10.0.2.1/24
+!
+interface eth3
+ ip address 10.0.3.1/24
+!
diff --git a/tests/mutests/mgre_four_peers/u0/etc.frr/vtysh.conf b/tests/mutests/mgre_four_peers/u0/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_four_peers/u0/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/README.md b/tests/mutests/mgre_mixed_unicast_mcast/README.md
new file mode 100644
index 0000000..9debbb0
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/README.md
@@ -0,0 +1,52 @@
+# Mixed underlay unicast/multicast mGRE mutest
+
+Five routers (r0..r4) share one mGRE overlay (`remote 0.0.0.0`,
+`172.16.0.0/24`). The underlay is mixed: r0/r1/r2 can take GRE-in-
+multicast, r3/r4 cannot.
+
+```
+ h0 -- r0 -- lan0 --\ mcast-capable
+ h1 -- r1 -- lan1 ---\
+ h2 -- r2 -- lan2 ---- u0 rmerge eth0,eth1,eth2 only
+ h3 -- r3 -- lan3 ---/ unicast-only
+ h4 -- r4 -- lan4 --/
+```
+
+`u0` unicasts among all five LANs. Its nrlsmf `rmerge` is only on the
+three mcast-capable LANs (PIM stand-in). r3 and r4 are not in that
+flood, so they never see underlay group `239.1.1.1`.
+
+Every router uses the same wildcard-remote mGRE device (`gre1`). Overlay
+multicast inject:
+
+* r0/r1/r2: `map gre1,,239.1.1.1` plus explicit unicast maps to
+ r3 and r4, and `ujoin 239.1.1.1,eth0`
+* r3/r4: explicit unicast `map` to every other router (no `ujoin`)
+
+`gre1` is `layered` so overlay multicast that arrived on the tunnel is
+not sent back out `gre1`. Sources on `eth1` still inject onto `gre1`.
+
+Overlay unicast uses kernel `ip neigh`. Application multicast is sourced
+on `h0` and received on `h1`..`h4`. While that flow is running, the test
+captures 24 GRE packets leaving `r0` `eth0` (8 overlay pps × 3 inject
+dests) and checks they are 8 to `239.1.1.1`, 8 to r3, and 8 to r4.
+
+After that CF pass, the same overlay nrlsmf processes get runtime
+`with-frr` and `elastic overlay` (maps, `layered gre1`, and grouping
+stay). Two keeper cases:
+
+* Keep multicast-underlay neighbor `h1` at 8 pps; stop `h2`/`h3`/`h4`
+ (at most 1 pps).
+* Then only unicast-only neighbor `h3` joins: `h3` at 8 pps; the other
+ multicast neighbors (`h1`/`h2`) and the other unicast neighbor (`h4`)
+ at most 1 pps.
+
+FRR `pimd` serves IGMP on each router's host `eth1`.
+
+## Run
+
+From `tests/mutests` (requires root, `nrlsmf` on PATH):
+
+```bash
+sudo mutest mgre_mixed_unicast_mcast
+```
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/mixed_hosts.py b/tests/mutests/mgre_mixed_unicast_mcast/mixed_hosts.py
new file mode 100644
index 0000000..2782e56
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/mixed_hosts.py
@@ -0,0 +1,155 @@
+"""Shared host-LAN helpers for the mixed unicast/mcast mGRE mutest.
+
+Each router has an application host off eth1. Iperf is sourced on h0
+(off r0) and received on h1/h2/h3/h4 so nrlsmf CF is both the first
+hop onto the overlay and the last hop off it.
+"""
+
+import time
+
+HOST_IFACE = "eth0"
+ROUTER_HOST_IFACE = "eth1"
+IPERF_TTL = "16"
+
+# (router, router eth1 addr, host name, host addr)
+HOST_LANS = (
+ ("r0", "192.168.55.1", "h0", "192.168.55.2"),
+ ("r1", "192.168.56.1", "h1", "192.168.56.2"),
+ ("r2", "192.168.57.1", "h2", "192.168.57.2"),
+ ("r3", "192.168.58.1", "h3", "192.168.58.2"),
+ ("r4", "192.168.59.1", "h4", "192.168.59.2"),
+)
+
+SOURCE_HOST = "h0"
+SOURCE_HOST_ADDR = "192.168.55.2"
+RECV_HOSTS = ("h1", "h2", "h3", "h4")
+
+
+def setup_host_lan(step, wait_step):
+ """Bring up each router's host LAN and its application host."""
+ for router, router_addr, host, host_addr in HOST_LANS:
+ step(router, f"ethtool -K {ROUTER_HOST_IFACE} rx off tx off || true")
+ wait_step(
+ router,
+ f"ip -br addr show dev {ROUTER_HOST_IFACE}",
+ match=router_addr,
+ desc=f"{router} {ROUTER_HOST_IFACE} address {router_addr}",
+ timeout=30,
+ )
+ step(router, "sysctl -w net.ipv4.conf.all.mc_forwarding=0 || true")
+
+ step(host, f"ethtool -K {HOST_IFACE} rx off tx off || true")
+ step(host, f"ip addr add {host_addr}/24 dev {HOST_IFACE} || true")
+ step(host, f"ip link set {HOST_IFACE} up")
+ step(host, f"ip route replace 0.0.0.0/0 via {router_addr}")
+ wait_step(
+ host,
+ f"ip -br addr show dev {HOST_IFACE}",
+ match=host_addr,
+ desc=f"{host} {HOST_IFACE} address {host_addr}",
+ timeout=30,
+ )
+ wait_step(
+ host,
+ f"ping -c1 -W2 {router_addr}",
+ match="1 received",
+ desc=f"{host} reaches {router} on the host LAN",
+ timeout=20,
+ )
+
+
+def start_overlay_mcast_servers(step, receivers, mcast):
+ for name in receivers:
+ step(name, f"ip route replace {mcast}/32 dev {HOST_IFACE}")
+ step(
+ name,
+ f"iperf -u -T 4 -i 1 -s -e -B {mcast}%{HOST_IFACE} "
+ f"> iperf-{name}-server.log 2>&1 &",
+ )
+
+
+def restart_overlay_mcast_servers(step, receivers, mcast):
+ """New iperf server log. Do not truncate a live iperf file."""
+ for name in receivers:
+ step(name, "pkill iperf || true")
+ time.sleep(1)
+ start_overlay_mcast_servers(step, receivers, mcast)
+
+
+def start_host_mcast_client(step, wait_step, mcast):
+ step(SOURCE_HOST, f"ip route replace {mcast}/32 dev {HOST_IFACE}")
+ step(
+ SOURCE_HOST,
+ f"iperf -u -T {IPERF_TTL} -t 1000 -i 1 -b 8pps -l 1024 -e "
+ f"-B {SOURCE_HOST_ADDR} -c {mcast} &> iperf-{SOURCE_HOST}-client.log &",
+ )
+ wait_step(
+ SOURCE_HOST,
+ f"tail -n1 iperf-{SOURCE_HOST}-client.log",
+ match="8 pps",
+ desc="h0 sending application multicast at 8 pps",
+ timeout=30,
+ )
+
+
+def wait_overlay_mcast_receivers(wait_step, receivers):
+ for name in receivers:
+ wait_step(
+ name,
+ f'grep "8 pps" iperf-{name}-server.log',
+ match="8 pps",
+ desc=f"{name} receiving application multicast at 8 pps",
+ timeout=20,
+ )
+
+
+def count_overlay_mcast_pkts(step, node, mcast, iface=HOST_IFACE, window_s=2):
+ """Count packets to ``mcast`` on ``iface`` over ``window_s`` seconds."""
+ raw = step(
+ node,
+ f"timeout {window_s} tcpdump -nn -l -i {iface} host {mcast} 2>/dev/null "
+ f"| grep -c {mcast} || true",
+ )
+ n = 0
+ for tok in str(raw).split():
+ if tok.isdigit():
+ n = int(tok)
+ return n
+
+
+def cleanup_iperf(step, receivers):
+ step(SOURCE_HOST, "pkill iperf || true")
+ for name in receivers:
+ step(name, "pkill iperf || true")
+
+
+def enable_host_igmp(step, wait_step, routers, host_iface=ROUTER_HOST_IFACE):
+ """Enable IGMP on each router's host LAN.
+
+ FRR serves IGMP from pimd. No ``ip pim`` on the iface.
+ """
+ for name in routers:
+ step(name, "pgrep -x pimd >/dev/null || /usr/lib/frr/pimd -d")
+ step(
+ name,
+ "vtysh -c 'configure terminal' "
+ f"-c 'interface {host_iface}' "
+ "-c 'ip igmp'",
+ )
+ wait_step(
+ name,
+ f"vtysh -c 'show ip igmp interface {host_iface}'",
+ match=host_iface,
+ desc=f"{name} FRR IGMP enabled on {host_iface}",
+ timeout=20,
+ )
+
+
+def wait_igmp_group(wait_step, router, mcast, timeout=30):
+ wait_step(
+ router,
+ "vtysh -c 'show ip igmp groups'",
+ match=mcast,
+ desc=f"{router} FRR IGMP has {mcast}",
+ timeout=timeout,
+ )
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/munet.yaml b/tests/mutests/mgre_mixed_unicast_mcast/munet.yaml
new file mode 100644
index 0000000..a44f0a2
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/munet.yaml
@@ -0,0 +1,77 @@
+# Five GRE/mGRE routers with an underlay router between the LANs,
+# plus an application host off each SMF router.
+#
+# h0 -- r0 -- lan0 --\ mcast-capable (u0 rmerge)
+# h1 -- r1 -- lan1 ---\
+# h2 -- r2 -- lan2 ---- u0
+# h3 -- r3 -- lan3 ---/ unicast-only (no underlay mcast)
+# h4 -- r4 -- lan4 --/
+#
+# u0 routes unicast among all five LANs. Underlay multicast is only
+# flooded among lan0..lan2 (r0/r1/r2). r3 and r4 stay unicast-only.
+topology:
+ networks:
+ - name: lan0
+ - name: lan1
+ - name: lan2
+ - name: lan3
+ - name: lan4
+ - name: lanh0
+ - name: lanh1
+ - name: lanh2
+ - name: lanh3
+ - name: lanh4
+ nodes:
+ - name: u0
+ kind: frr
+ connections:
+ - to: lan0
+ - to: lan1
+ - to: lan2
+ - to: lan3
+ - to: lan4
+ - name: h0
+ kind: host
+ connections:
+ - to: lanh0
+ - name: h1
+ kind: host
+ connections:
+ - to: lanh1
+ - name: h2
+ kind: host
+ connections:
+ - to: lanh2
+ - name: h3
+ kind: host
+ connections:
+ - to: lanh3
+ - name: h4
+ kind: host
+ connections:
+ - to: lanh4
+ - name: r0
+ kind: frr
+ connections:
+ - to: lan0
+ - to: lanh0
+ - name: r1
+ kind: frr
+ connections:
+ - to: lan1
+ - to: lanh1
+ - name: r2
+ kind: frr
+ connections:
+ - to: lan2
+ - to: lanh2
+ - name: r3
+ kind: frr
+ connections:
+ - to: lan3
+ - to: lanh3
+ - name: r4
+ kind: frr
+ connections:
+ - to: lan4
+ - to: lanh4
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/mutest_mgre_mixed.py b/tests/mutests/mgre_mixed_unicast_mcast/mutest_mgre_mixed.py
new file mode 100644
index 0000000..8d9145f
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/mutest_mgre_mixed.py
@@ -0,0 +1,468 @@
+"""Example: mGRE with mixed underlay multicast and unicast peers.
+
+Topology:
+
+ h0 -- r0 -- lan0 --\\ mcast-capable
+ h1 -- r1 -- lan1 ---\\
+ h2 -- r2 -- lan2 ---- u0 (rmerge eth0,eth1,eth2 only)
+ h3 -- r3 -- lan3 ---/ unicast-only
+ h4 -- r4 -- lan4 --/
+
+All five routers share one wildcard-remote mGRE interface (gre1,
+remote 0.0.0.0) on 172.16.0.0/24. Overlay unicast uses kernel
+ip neigh. Overlay multicast inject is mixed:
+
+* r0/r1/r2 join underlay group 239.1.1.1 (ujoin on eth0) and map that
+ group as one inject dest, plus explicit unicast maps to r3 and r4.
+* r3/r4 do not join the group; they map every other router as a
+ unicast inject dest.
+
+gre1 is layered so a packet received on the tunnel is not flooded
+back out gre1 (unicast-only routers do not re-inject to the mcast
+set). Traffic that arrives on eth1 is still injected onto gre1.
+The overlay-multicast section captures GRE leaving r0 eth0 and
+checks 8 overlay pps become 24 GRE (8 to 239.1.1.1, 8 to r3, 8 to r4).
+
+After that CF pass, the same overlay nrlsmf processes get runtime
+``with-frr`` and ``elastic overlay``. Two keeper cases:
+
+* Multicast-underlay neighbor h1 at 8 pps; stop h2/h3/h4 (at most 1 pps).
+* Then unicast-only neighbor h3 at 8 pps; stop h1/h2/h4 (at most 1 pps).
+
+u0 unicasts among all five LANs. Its nrlsmf rmerge is only on
+eth0,eth1,eth2 so underlay multicast never reaches r3/r4. That nrlsmf
+instance is a PIM stand-in only; it is not part of the overlay.
+"""
+
+from datetime import datetime
+from datetime import timedelta
+
+from munet.mutest.userapi import get_target
+from munet.mutest.userapi import script_dir
+from munet.mutest.userapi import section
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+from munet.mutest.userapi import wait_step
+
+import sys
+import time
+
+sys.path.insert(0, str(script_dir()))
+sys.path.insert(0, str(script_dir().parent))
+from mixed_hosts import RECV_HOSTS
+from mixed_hosts import ROUTER_HOST_IFACE
+from mixed_hosts import SOURCE_HOST
+from mixed_hosts import cleanup_iperf
+from mixed_hosts import count_overlay_mcast_pkts
+from mixed_hosts import enable_host_igmp
+from mixed_hosts import restart_overlay_mcast_servers
+from mixed_hosts import setup_host_lan
+from mixed_hosts import start_host_mcast_client
+from mixed_hosts import start_overlay_mcast_servers
+from mixed_hosts import wait_igmp_group
+from mixed_hosts import wait_overlay_mcast_receivers
+from smf_cli import check_common_show
+from smf_cli import check_group_elastic
+from smf_cli import check_show_neighbors
+from smf_cli import check_show_tunnel
+from smf_cli import cli_cmd
+from smf_cli import wait_em_flow
+from smf_cli import wait_em_managed
+
+ROUTERS = {
+ "r0": {"underlay": "10.0.0.2", "overlay": "172.16.0.1"},
+ "r1": {"underlay": "10.0.1.2", "overlay": "172.16.0.2"},
+ "r2": {"underlay": "10.0.2.2", "overlay": "172.16.0.3"},
+ "r3": {"underlay": "10.0.3.2", "overlay": "172.16.0.4"},
+ "r4": {"underlay": "10.0.4.2", "overlay": "172.16.0.5"},
+}
+
+MCAST_ROUTERS = ("r0", "r1", "r2")
+UCAST_ONLY = ("r3", "r4")
+
+UNDERLAY_MCAST = "239.1.1.1"
+OVERLAY_MCAST = "239.0.0.1"
+GRE_DEV = "gre1"
+U0_MCAST_IFACES = "eth0,eth1,eth2"
+GRE_DUMP = "tcpdump-r0-eth0-gre.log"
+GRE_WINDOW_S = 1.0
+GRE_PPS = 8
+GRE_PPS_SLOP = 2
+
+
+def gre_dest_counts(path, window_s=GRE_WINDOW_S):
+ """Count GRE dests in the first ``window_s`` seconds of a tcpdump log."""
+ samples = []
+ for ln in path.read_text().splitlines():
+ if "GREv0" not in ln:
+ continue
+ parts = ln.split()
+ try:
+ ts = datetime.strptime(parts[0], "%H:%M:%S.%f")
+ except (ValueError, IndexError):
+ continue
+ dest = parts[4].rstrip(":")
+ samples.append((ts, dest))
+ counts = {}
+ if not samples:
+ return counts
+ t1 = samples[0][0] + timedelta(seconds=window_s)
+ for ts, dest in samples:
+ if ts >= t1:
+ break
+ counts[dest] = counts.get(dest, 0) + 1
+ return counts
+
+
+def gre_pps_ok(n):
+ return GRE_PPS - GRE_PPS_SLOP <= n <= GRE_PPS + GRE_PPS_SLOP
+
+
+section("Disable offloads and wait for underlay addresses")
+
+step("u0", "sysctl -w net.ipv4.ip_forward=1")
+step("u0", "sysctl -w net.ipv4.conf.all.rp_filter=0")
+step("u0", "sysctl -w net.ipv4.conf.default.rp_filter=0")
+step("u0", "sysctl -w net.ipv4.conf.all.send_redirects=0")
+
+for name, cfg in ROUTERS.items():
+ step(name, "ethtool -K eth0 rx off tx off || true")
+ step(name, "sysctl -w net.ipv4.conf.all.rp_filter=0")
+ step(name, "sysctl -w net.ipv4.conf.eth0.rp_filter=0")
+ step(name, "sysctl -w net.ipv4.conf.all.send_redirects=0")
+ wait_step(
+ name,
+ "ip -br addr show dev eth0",
+ match=cfg["underlay"],
+ desc=f"{name} underlay address {cfg['underlay']}",
+ timeout=30,
+ )
+
+for ifname, addr in (
+ ("eth0", "10.0.0.1"),
+ ("eth1", "10.0.1.1"),
+ ("eth2", "10.0.2.1"),
+ ("eth3", "10.0.3.1"),
+ ("eth4", "10.0.4.1"),
+):
+ step("u0", f"ethtool -K {ifname} rx off tx off || true")
+ step("u0", f"sysctl -w net.ipv4.conf.{ifname}.rp_filter=0")
+ wait_step(
+ "u0",
+ f"ip -br addr show dev {ifname}",
+ match=addr,
+ desc=f"u0 {ifname} address {addr}",
+ timeout=30,
+ )
+
+setup_host_lan(step, wait_step)
+
+section("Underlay unicast reachability through u0")
+
+for src in ROUTERS:
+ for dst, dcfg in ROUTERS.items():
+ if src == dst:
+ continue
+ wait_step(
+ src,
+ f"ping -c1 -W2 {dcfg['underlay']}",
+ match="1 received",
+ desc=f"{src} ping underlay {dst} ({dcfg['underlay']})",
+ timeout=20,
+ )
+
+section("Underlay multicast relay on u0 (r0/r1/r2 only)")
+
+# PIM stand-in for the mcast-capable LANs only. eth3/eth4 (r3/r4) are
+# left out so those routers never see UNDERLAY_MCAST.
+step(
+ "u0",
+ "nrlsmf debug 4 "
+ "instance smf-u0-underlay "
+ f"rmerge {U0_MCAST_IFACES} "
+ "&> nrlsmf-u0-underlay.log &",
+)
+wait_step(
+ "u0",
+ 'pgrep -af "nrlsmf.*instance smf-u0-underlay"',
+ match="smf-u0-underlay",
+ desc="u0 underlay nrlsmf running",
+ timeout=20,
+)
+wait_step(
+ "u0",
+ 'grep "regular group" nrlsmf-u0-underlay.log',
+ match="merge",
+ desc="u0 nrlsmf log shows merge group",
+ timeout=20,
+)
+
+section("Create wildcard-remote mGRE tunnels (remote 0.0.0.0)")
+
+for name, cfg in ROUTERS.items():
+ step(name, "ip addr flush dev gre0 2>/dev/null || true")
+ step(name, "ip link set gre0 down 2>/dev/null || true")
+ step(name, f"ip link del {GRE_DEV} 2>/dev/null || true")
+ step(
+ name,
+ f"ip link add name {GRE_DEV} type gre "
+ f"local {cfg['underlay']} remote 0.0.0.0 ttl 64",
+ )
+ step(name, f"ip addr add {cfg['overlay']}/24 dev {GRE_DEV}")
+ step(name, f"ip link set {GRE_DEV} multicast on")
+ step(name, f"ip link set {GRE_DEV} up")
+ for other, ocfg in ROUTERS.items():
+ if other == name:
+ continue
+ step(
+ name,
+ f"ip neigh replace {ocfg['overlay']} lladdr {ocfg['underlay']} "
+ f"nud permanent dev {GRE_DEV}",
+ )
+ wait_step(
+ name,
+ f"ip -br link show {GRE_DEV}",
+ match="UP",
+ desc=f"{name} {GRE_DEV} is UP",
+ )
+
+for name in MCAST_ROUTERS:
+ step(name, f"ip route replace {UNDERLAY_MCAST}/32 dev eth0")
+
+section("Overlay unicast across mGRE (before overlay nrlsmf)")
+
+for src, scfg in ROUTERS.items():
+ for dst, dcfg in ROUTERS.items():
+ if src == dst:
+ continue
+ wait_step(
+ src,
+ f"ping -c1 -W3 -I {scfg['overlay']} {dcfg['overlay']}",
+ match="1 received",
+ desc=f"{src} ping overlay {dst} ({dcfg['overlay']})",
+ timeout=20,
+ )
+
+section("Start overlay nrlsmf (mixed inject dests)")
+
+for name, cfg in ROUTERS.items():
+ if name in MCAST_ROUTERS:
+ maps = (
+ f"map {GRE_DEV},{cfg['underlay']},{UNDERLAY_MCAST} "
+ + " ".join(
+ f"map {GRE_DEV},{cfg['underlay']},{ROUTERS[peer]['underlay']}"
+ for peer in UCAST_ONLY
+ )
+ )
+ ujoin = f"ujoin {UNDERLAY_MCAST},eth0 "
+ else:
+ maps = " ".join(
+ f"map {GRE_DEV},{cfg['underlay']},{ocfg['underlay']}"
+ for other, ocfg in ROUTERS.items()
+ if other != name
+ )
+ ujoin = ""
+ step(
+ name,
+ "nrlsmf debug 4 "
+ f"instance smf-{name}-mixed "
+ f"add overlay,cf,eth1,{GRE_DEV} "
+ f"layered {GRE_DEV} "
+ f"{ujoin}"
+ f"{maps} "
+ "&> nrlsmf-mgre-mixed.log &",
+ )
+ wait_step(
+ name,
+ f'pgrep -af "nrlsmf.*instance smf-{name}-mixed"',
+ match=f"smf-{name}-mixed",
+ desc=f"{name} nrlsmf running on {GRE_DEV}",
+ timeout=20,
+ )
+ wait_step(
+ name,
+ 'grep "regular group" nrlsmf-mgre-mixed.log',
+ match="overlay",
+ desc=f"{name} nrlsmf log shows overlay group",
+ timeout=20,
+ )
+
+section("nrlsmf --cli show tunnel / neighbors (json, mixed inject dests)")
+
+# r0/r1/r2 map underlay mcast + unicast-only remotes; r3/r4 map every other
+# router as unicast. Overlay pings populate Neighbor IP from kernel neigh.
+for name, cfg in ROUTERS.items():
+ inst = f"smf-{name}-mixed"
+ check_common_show(name, inst, group_name="overlay", ifaces=("eth1", GRE_DEV))
+ if name in MCAST_ROUTERS:
+ remotes = [UNDERLAY_MCAST] + [ROUTERS[p]["underlay"] for p in UCAST_ONLY]
+ else:
+ remotes = [ocfg["underlay"] for other, ocfg in ROUTERS.items() if other != name]
+ check_show_tunnel(
+ name, inst, GRE_DEV,
+ local=cfg["underlay"],
+ remotes=remotes,
+ overlay_ip=cfg["overlay"],
+ want_c=True,
+ )
+ check_show_neighbors(
+ name, inst, GRE_DEV,
+ remotes=remotes,
+ neighbor_ips=[ocfg["overlay"] for other, ocfg in ROUTERS.items() if other != name],
+ min_count=len(remotes),
+ want_c=True,
+ )
+
+section("[Mixed] Overlay multicast: h0 -> SMF -> h1/h2/h3/h4")
+
+RECEIVERS = RECV_HOSTS
+start_overlay_mcast_servers(step, RECEIVERS, OVERLAY_MCAST)
+start_host_mcast_client(step, wait_step, OVERLAY_MCAST)
+
+# 8 overlay pps x 3 mapped remotes = 24 GRE/s on r0 eth0. Capture a
+# few seconds into r0's rundir, then count dests in a 1s timestamp
+# window (not the first N packets).
+step(
+ "r0",
+ "timeout 3 tcpdump -lnni eth0 "
+ f"'proto gre and src host {ROUTERS['r0']['underlay']}' "
+ f"&> {GRE_DUMP} || true",
+)
+counts = gre_dest_counts(get_target("r0").rundir / GRE_DUMP)
+n_mcast = counts.get(UNDERLAY_MCAST, 0)
+n_r3 = counts.get(ROUTERS["r3"]["underlay"], 0)
+n_r4 = counts.get(ROUTERS["r4"]["underlay"], 0)
+n_r1 = counts.get(ROUTERS["r1"]["underlay"], 0)
+n_r2 = counts.get(ROUTERS["r2"]["underlay"], 0)
+test_step(
+ gre_pps_ok(n_mcast)
+ and gre_pps_ok(n_r3)
+ and gre_pps_ok(n_r4)
+ and n_r1 == 0
+ and n_r2 == 0,
+ f"r0 eth0 1s GRE: mcast={n_mcast} r3={n_r3} r4={n_r4} "
+ f"r1={n_r1} r2={n_r2} (expect {GRE_PPS}±{GRE_PPS_SLOP} each)",
+ "r0",
+)
+
+wait_overlay_mcast_receivers(wait_step, RECEIVERS)
+
+KEEP_RECV_MCAST = ("h1",)
+IDLE_RECV_MCAST = ("h2", "h3", "h4")
+KEEP_RECV_UCAST = ("h3",)
+IDLE_RECV_UCAST = ("h1", "h2", "h4")
+EM_WINDOW_S = 2.0
+EM_IDLE_MAX_PPS = 1
+
+
+section("[Mixed] Host IGMP on eth1; h0 joins 239.0.0.1")
+
+# h0 is the sender; join 239.0.0.1 the same way h1..h4 do so r0
+# eth1 is an IGMP host LAN and with-frr sees a local member.
+start_overlay_mcast_servers(step, (SOURCE_HOST,), OVERLAY_MCAST)
+enable_host_igmp(step, wait_step, ROUTERS)
+for name in ROUTERS:
+ wait_igmp_group(wait_step, name, OVERLAY_MCAST)
+
+section("[Mixed] Runtime with-frr + elastic overlay")
+
+# Same running CF processes. Maps, layered gre1, and overlay grouping
+# stay. with-frr starts FRR IGMP polling; elastic overlay turns on EM.
+for name in ROUTERS:
+ inst = f"smf-{name}-mixed"
+ cli_cmd(name, "with-frr", inst)
+ cli_cmd(name, "elastic overlay", inst)
+ check_group_elastic(
+ name, inst, "overlay", enabled=True,
+ ifaces=(ROUTER_HOST_IFACE, GRE_DEV),
+ )
+
+section("[Mixed] EM ack/forwarding state")
+
+# r1 has a local IGMP member (h1): it must ACK upstream so r0 FORWARDs
+# on gre1. Default FwdStatus is usually LIMIT; oif is the per-iface
+# token-bucket status after EM_ACK (show groups details).
+wait_em_flow("r1", "smf-r1-mixed", OVERLAY_MCAST, ack="yes", status="ACTIVE", timeout=20)
+wait_em_flow("r0", "smf-r0-mixed", OVERLAY_MCAST, ack="yes", timeout=20)
+wait_em_managed("r1", "smf-r1-mixed", ROUTER_HOST_IFACE, OVERLAY_MCAST, timeout=20)
+
+section("[Mixed] Overlay mcast after elastic")
+
+# New h1 iperf log so grep cannot match CF-era lines.
+restart_overlay_mcast_servers(step, KEEP_RECV_MCAST, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, KEEP_RECV_MCAST)
+
+section("[Mixed] Stop h2/h3/h4; keep multicast neighbor h1")
+
+for name in IDLE_RECV_MCAST:
+ step(name, "pkill iperf || true")
+
+time.sleep(6)
+wait_step(
+ "h1",
+ "tail -n1 iperf-h1-server.log",
+ match="8 pps",
+ desc="h1 still receiving application multicast at 8 pps after elastic",
+ timeout=30,
+)
+
+for name in IDLE_RECV_MCAST:
+ n = count_overlay_mcast_pkts(step, name, OVERLAY_MCAST, window_s=EM_WINDOW_S)
+ pps = n / EM_WINDOW_S
+ test_step(
+ pps <= EM_IDLE_MAX_PPS,
+ f"{name} non-receiver overlay mcast {pps:.1f} pps (at most {EM_IDLE_MAX_PPS})",
+ target=name,
+ )
+
+section("[Mixed] Unicast neighbor h3 joins; rate-limit mcast peers and r4")
+
+# r3 is a unicast inject dest. After h1 leaves, only r3 should ACK so
+# r0 FORWARDs the r3 map and LIMITs underlay 239.1.1.1 and the r4 map.
+for name in IDLE_RECV_UCAST:
+ step(name, "pkill iperf || true")
+restart_overlay_mcast_servers(step, KEEP_RECV_UCAST, OVERLAY_MCAST)
+wait_igmp_group(wait_step, "r3", OVERLAY_MCAST)
+wait_em_flow(
+ "r3", "smf-r3-mixed", OVERLAY_MCAST, ack="yes", status="ACTIVE", timeout=20,
+)
+wait_em_flow("r0", "smf-r0-mixed", OVERLAY_MCAST, ack="yes", timeout=20)
+wait_em_managed("r3", "smf-r3-mixed", ROUTER_HOST_IFACE, OVERLAY_MCAST, timeout=20)
+
+restart_overlay_mcast_servers(step, KEEP_RECV_UCAST, OVERLAY_MCAST)
+wait_overlay_mcast_receivers(wait_step, KEEP_RECV_UCAST)
+
+time.sleep(6)
+wait_step(
+ "h3",
+ "tail -n1 iperf-h3-server.log",
+ match="8 pps",
+ desc="h3 receiving application multicast at 8 pps (unicast-only last hop)",
+ timeout=30,
+)
+
+for name in IDLE_RECV_UCAST:
+ n = count_overlay_mcast_pkts(step, name, OVERLAY_MCAST, window_s=EM_WINDOW_S)
+ pps = n / EM_WINDOW_S
+ test_step(
+ pps <= EM_IDLE_MAX_PPS,
+ f"{name} non-receiver overlay mcast {pps:.1f} pps (at most {EM_IDLE_MAX_PPS})",
+ target=name,
+ )
+
+section("Cleanup")
+
+cleanup_iperf(step, RECEIVERS)
+step("r0", "pkill tcpdump || true")
+for name in ROUTERS:
+ step(name, "pkill nrlsmf || true")
+ wait_step(
+ name,
+ "pgrep -af nrlsmf || true",
+ match="",
+ desc=f"{name} nrlsmf stopped",
+ timeout=15,
+ )
+
+step("u0", "pkill nrlsmf || true")
+test_step(True, "mGRE mixed unicast/mcast five-router mutest completed")
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/daemons b/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/daemons
new file mode 100644
index 0000000..6d87a20
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/frr.conf b/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/frr.conf
new file mode 100644
index 0000000..8b72b6e
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/frr.conf
@@ -0,0 +1,14 @@
+log file /var/log/frr/frr.log
+!
+! r0: SMF router, underlay-multicast capable. eth0 toward u0; eth1
+! toward h0 (iperf source).
+!
+ip route 0.0.0.0/0 10.0.0.1
+!
+interface eth0
+ ip address 10.0.0.2/24
+!
+interface eth1
+ ip address 192.168.55.1/24
+ ip igmp
+!
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/vtysh.conf b/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r0/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/daemons b/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/daemons
new file mode 100644
index 0000000..6d87a20
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/frr.conf b/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/frr.conf
new file mode 100644
index 0000000..88ff7ea
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/frr.conf
@@ -0,0 +1,14 @@
+log file /var/log/frr/frr.log
+!
+! r1: SMF router, underlay-multicast capable. eth0 toward u0; eth1
+! toward h1 (iperf receiver).
+!
+ip route 0.0.0.0/0 10.0.1.1
+!
+interface eth0
+ ip address 10.0.1.2/24
+!
+interface eth1
+ ip address 192.168.56.1/24
+ ip igmp
+!
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/vtysh.conf b/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r1/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/daemons b/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/daemons
new file mode 100644
index 0000000..6d87a20
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/frr.conf b/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/frr.conf
new file mode 100644
index 0000000..8927c24
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/frr.conf
@@ -0,0 +1,14 @@
+log file /var/log/frr/frr.log
+!
+! r2: SMF router, underlay-multicast capable. eth0 toward u0; eth1
+! toward h2 (iperf receiver).
+!
+ip route 0.0.0.0/0 10.0.2.1
+!
+interface eth0
+ ip address 10.0.2.2/24
+!
+interface eth1
+ ip address 192.168.57.1/24
+ ip igmp
+!
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/vtysh.conf b/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r2/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/daemons b/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/daemons
new file mode 100644
index 0000000..6d87a20
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/frr.conf b/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/frr.conf
new file mode 100644
index 0000000..9a56f82
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/frr.conf
@@ -0,0 +1,14 @@
+log file /var/log/frr/frr.log
+!
+! r3: SMF router, underlay unicast-only. eth0 toward u0; eth1 toward
+! h3 (iperf receiver).
+!
+ip route 0.0.0.0/0 10.0.3.1
+!
+interface eth0
+ ip address 10.0.3.2/24
+!
+interface eth1
+ ip address 192.168.58.1/24
+ ip igmp
+!
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/vtysh.conf b/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r3/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/daemons b/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/daemons
new file mode 100644
index 0000000..6d87a20
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/daemons
@@ -0,0 +1,2 @@
+nhrpd=yes
+pimd=yes
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/frr.conf b/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/frr.conf
new file mode 100644
index 0000000..19a06ab
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/frr.conf
@@ -0,0 +1,14 @@
+log file /var/log/frr/frr.log
+!
+! r4: SMF router, underlay unicast-only. eth0 toward u0; eth1 toward
+! h4 (iperf receiver).
+!
+ip route 0.0.0.0/0 10.0.4.1
+!
+interface eth0
+ ip address 10.0.4.2/24
+!
+interface eth1
+ ip address 192.168.59.1/24
+ ip igmp
+!
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/vtysh.conf b/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/r4/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/u0/etc.frr/daemons b/tests/mutests/mgre_mixed_unicast_mcast/u0/etc.frr/daemons
new file mode 100644
index 0000000..e69de29
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/u0/etc.frr/frr.conf b/tests/mutests/mgre_mixed_unicast_mcast/u0/etc.frr/frr.conf
new file mode 100644
index 0000000..9b9c4d0
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/u0/etc.frr/frr.conf
@@ -0,0 +1,21 @@
+log file /var/log/frr/frr.log
+!
+! Underlay ("u0"): unicast among all five LANs. nrlsmf rmerge on
+! eth0..eth2 only (see mutest) so r0/r1/r2 see underlay multicast
+! and r3/r4 do not.
+!
+interface eth0
+ ip address 10.0.0.1/24
+!
+interface eth1
+ ip address 10.0.1.1/24
+!
+interface eth2
+ ip address 10.0.2.1/24
+!
+interface eth3
+ ip address 10.0.3.1/24
+!
+interface eth4
+ ip address 10.0.4.1/24
+!
diff --git a/tests/mutests/mgre_mixed_unicast_mcast/u0/etc.frr/vtysh.conf b/tests/mutests/mgre_mixed_unicast_mcast/u0/etc.frr/vtysh.conf
new file mode 100644
index 0000000..e0ab9cb
--- /dev/null
+++ b/tests/mutests/mgre_mixed_unicast_mcast/u0/etc.frr/vtysh.conf
@@ -0,0 +1 @@
+service integrated-vtysh-config
diff --git a/tests/mutests/smf_cli.py b/tests/mutests/smf_cli.py
new file mode 100644
index 0000000..39f5fe4
--- /dev/null
+++ b/tests/mutests/smf_cli.py
@@ -0,0 +1,300 @@
+"""Helpers to query a running nrlsmf via ``nrlsmf --cli`` JSON show commands."""
+
+import json
+import time
+
+from munet.mutest.userapi import step
+from munet.mutest.userapi import test_step
+
+
+def cli_cmd(node, command, instance=None):
+ inst = f"-i {instance} " if instance else ""
+ return step(node, f"nrlsmf --cli {inst}-c {json.dumps(command)}")
+
+
+def parse_json_blob(text):
+ text = text.strip()
+ start = None
+ for i, ch in enumerate(text):
+ if ch in "[{":
+ start = i
+ break
+ if start is None:
+ raise ValueError(f"no JSON object in {text[:200]!r}")
+ text = text[start:]
+ end_obj = text.rfind("}")
+ end_arr = text.rfind("]")
+ end = max(end_obj, end_arr)
+ if end >= 0:
+ text = text[: end + 1]
+ return json.loads(text)
+
+
+def show_json(node, show_cmd, instance=None):
+ raw = cli_cmd(node, f"{show_cmd} json", instance)
+ try:
+ data = parse_json_blob(raw)
+ except (json.JSONDecodeError, ValueError) as err:
+ test_step(
+ False,
+ f"{node} '{show_cmd} json' parse failed: {err}: {raw[:240]!r}",
+ target=node,
+ )
+ return None
+ return data
+
+
+def check_common_show(node, instance=None, group_name=None, ifaces=None):
+ """Ping plus JSON show version/statistics/interface/grouping."""
+ pong = cli_cmd(node, "ping", instance)
+ test_step("pong" in pong, f"{node} --cli ping returns pong", target=node)
+
+ ver = show_json(node, "show version", instance)
+ if ver is not None:
+ test_step(
+ isinstance(ver, dict) and bool(ver.get("jsonVersion")),
+ f"{node} show version json has jsonVersion",
+ target=node,
+ )
+
+ stats = show_json(node, "show statistics", instance)
+ if isinstance(stats, list):
+ names = {row.get("interface") for row in stats if isinstance(row, dict)}
+ for iface in ifaces or ():
+ test_step(iface in names, f"{node} statistics includes {iface}", target=node)
+
+ listing = show_json(node, "show interface", instance)
+ if isinstance(listing, list):
+ names = {row.get("Interface") for row in listing if isinstance(row, dict)}
+ for iface in ifaces or ():
+ test_step(iface in names, f"{node} interface list includes {iface}", target=node)
+
+ grouping = show_json(node, "show interface grouping", instance)
+ if isinstance(grouping, list) and group_name:
+ gnames = {g.get("GroupName") for g in grouping if isinstance(g, dict)}
+ test_step(group_name in gnames,
+ f"{node} grouping includes {group_name}", target=node)
+
+
+def check_group_elastic(node, instance, group_name, enabled, ifaces=None,
+ group_ifaces=None, absent_ifaces=None):
+ """Assert show interface grouping/interface report Elastic on or off."""
+ grouping = show_json(node, "show interface grouping", instance)
+ if grouping is None:
+ return
+ groups = [g for g in grouping if isinstance(g, dict) and g.get("GroupName") == group_name]
+ test_step(bool(groups), f"{node} grouping includes {group_name}", target=node)
+ if groups:
+ is_em = groups[0].get("Elastic") is True
+ test_step(
+ is_em is enabled,
+ f"{node} grouping {group_name} Elastic is {enabled}",
+ target=node,
+ )
+ have = set(groups[0].get("Interfaces") or [])
+ for iface in group_ifaces or ():
+ test_step(
+ iface in have,
+ f"{node} grouping {group_name} includes {iface}",
+ target=node,
+ )
+ for iface in absent_ifaces or ():
+ test_step(
+ iface not in have,
+ f"{node} grouping {group_name} omits {iface}",
+ target=node,
+ )
+ listing = show_json(node, "show interface", instance)
+ if listing is None or not ifaces:
+ return
+ want = "Elastic" if enabled else "Flood"
+ for iface in ifaces:
+ rows = [r for r in listing if isinstance(r, dict) and r.get("Interface") == iface]
+ test_step(bool(rows), f"{node} interface list includes {iface}", target=node)
+ if rows:
+ test_step(
+ rows[0].get("FwdMethod") == want,
+ f"{node} {iface} FwdMethod is {want}",
+ target=node,
+ )
+
+
+def check_show_groups(node, instance=None, mcast_addr=None):
+ groups = show_json(node, "show groups", instance)
+ if groups is None or not isinstance(groups, list):
+ return
+ if mcast_addr:
+ addrs = {g.get("MCastAddr") for g in groups if isinstance(g, dict)}
+ test_step(mcast_addr in addrs,
+ f"{node} show groups includes {mcast_addr}", target=node)
+
+
+def _em_flow(groups, mcast_addr):
+ if not isinstance(groups, list):
+ return None
+ for row in groups:
+ if isinstance(row, dict) and row.get("MCastAddr") == mcast_addr:
+ return row
+ return None
+
+
+def check_em_flow(node, instance, mcast_addr, ack=None, status=None, fwd_status=None):
+ """Assert show groups fields for an EM flow (Ack / Status / FwdStatus)."""
+ groups = show_json(node, "show groups", instance)
+ flow = _em_flow(groups, mcast_addr)
+ test_step(
+ flow is not None,
+ f"{node} show groups includes {mcast_addr}",
+ target=node,
+ )
+ if flow is None:
+ return None
+ if ack is not None:
+ have = flow.get("Ack")
+ test_step(
+ have == ack,
+ f"{node} {mcast_addr} Ack is {have!r} (want {ack!r})",
+ target=node,
+ )
+ if status is not None:
+ have = flow.get("Status")
+ test_step(
+ have == status,
+ f"{node} {mcast_addr} Status is {have!r} (want {status!r})",
+ target=node,
+ )
+ if fwd_status is not None:
+ have = flow.get("FwdStatus")
+ test_step(
+ have == fwd_status,
+ f"{node} {mcast_addr} FwdStatus is {have!r} (want {fwd_status!r})",
+ target=node,
+ )
+ return flow
+
+
+def _igmp_iface(rows, iface):
+ for row in rows or ():
+ if isinstance(row, dict) and row.get("Interface") == iface:
+ return row
+ return None
+
+
+def _igmp_has_group(rows, iface, group):
+ row = _igmp_iface(rows, iface)
+ if row is None or row.get("Managed") is not True:
+ return False
+ groups = row.get("Groups") if isinstance(row.get("Groups"), list) else []
+ return group in groups
+
+
+def wait_em_managed(node, instance, iface, group, timeout=20):
+ """Poll show igmp groups until iface is Managed and lists group."""
+ deadline = time.time() + timeout
+ last = None
+ while time.time() < deadline:
+ try:
+ last = parse_json_blob(cli_cmd(node, "show igmp groups json", instance))
+ except (json.JSONDecodeError, ValueError):
+ last = None
+ if _igmp_has_group(last, iface, group):
+ test_step(
+ True,
+ f"{node} show igmp groups {iface} has {group}",
+ target=node,
+ )
+ return last
+ time.sleep(1)
+ test_step(
+ False,
+ f"{node} show igmp groups {iface} missing {group}",
+ target=node,
+ )
+ return last
+
+
+def wait_em_flow(node, instance, mcast_addr, ack=None, status=None, timeout=20):
+ """Poll show groups until the flow exists and optional Ack/Status match."""
+ deadline = time.time() + timeout
+ last = None
+ while time.time() < deadline:
+ try:
+ last = parse_json_blob(cli_cmd(node, "show groups json", instance))
+ except (json.JSONDecodeError, ValueError):
+ last = None
+ flow = _em_flow(last, mcast_addr)
+ if flow is not None:
+ if ack is not None and flow.get("Ack") != ack:
+ time.sleep(1)
+ continue
+ if status is not None and flow.get("Status") != status:
+ time.sleep(1)
+ continue
+ return check_em_flow(
+ node, instance, mcast_addr, ack=ack, status=status,
+ )
+ time.sleep(1)
+ return check_em_flow(node, instance, mcast_addr, ack=ack, status=status)
+
+
+def wait_show_groups(node, instance, mcast_addr, timeout=20):
+ """Poll show groups until the EM FIB lists ``mcast_addr``."""
+ wait_em_flow(node, instance, mcast_addr, timeout=timeout)
+
+
+def _rows_for_iface(rows, iface):
+ if not isinstance(rows, list):
+ return []
+ return [r for r in rows if isinstance(r, dict) and r.get("Interface") == iface]
+
+
+def check_show_tunnel(node, instance, iface, local=None, remotes=None, overlay_ip=None,
+ want_c=None):
+ """Assert show tunnel json rows for iface (underlay Local/Remote, overlay IP)."""
+ rows = show_json(node, "show tunnel", instance)
+ if rows is None:
+ return
+ on_if = _rows_for_iface(rows, iface)
+ test_step(bool(on_if), f"{node} show tunnel has {iface}", target=node)
+ if local:
+ test_step(any(r.get("Local") == local for r in on_if),
+ f"{node} show tunnel Local {local} on {iface}", target=node)
+ if overlay_ip:
+ test_step(any(r.get("IP") == overlay_ip for r in on_if),
+ f"{node} show tunnel IP {overlay_ip} on {iface}", target=node)
+ if remotes:
+ have = {r.get("Remote") for r in on_if}
+ for remote in remotes:
+ test_step(remote in have,
+ f"{node} show tunnel Remote {remote} on {iface}", target=node)
+ if want_c is True:
+ test_step(any("C" in (r.get("Flags") or "") for r in on_if),
+ f"{node} show tunnel Flags include C on {iface}", target=node)
+ elif want_c is False:
+ test_step(all("C" not in (r.get("Flags") or "") for r in on_if),
+ f"{node} show tunnel Flags omit C on {iface}", target=node)
+
+
+def check_show_neighbors(node, instance, iface, remotes=None, neighbor_ips=None,
+ min_count=0, want_c=None):
+ """Assert show tunnel neighbors json for iface (Neighbor IP / Remote)."""
+ rows = show_json(node, "show tunnel neighbors", instance)
+ if rows is None:
+ return
+ on_if = _rows_for_iface(rows, iface)
+ test_step(len(on_if) >= min_count,
+ f"{node} show tunnel neighbors has >= {min_count} row(s) on {iface}",
+ target=node)
+ if remotes:
+ have = {r.get("Remote") for r in on_if}
+ for remote in remotes:
+ test_step(remote in have,
+ f"{node} neighbor Remote {remote} on {iface}", target=node)
+ if neighbor_ips:
+ have = {r.get("NeighborIP") for r in on_if}
+ for nip in neighbor_ips:
+ test_step(nip in have,
+ f"{node} Neighbor IP {nip} on {iface}", target=node)
+ if want_c is True:
+ test_step(any("C" in (r.get("Flags") or "") for r in on_if),
+ f"{node} neighbor Flags include C on {iface}", target=node)
|
|
|
|
|
|
|
|