ENGINEERING NOTE / LINUX NETWORKING

Why a route dump does not explain a Linux routing decision.

A route can exist, be more specific, and still not be the route Linux selects for the flow you are debugging.

THE TRAP

“The route is there” is not the same as “the kernel used it.”

Looking at ip route show table all is useful because it tells you what routes exist. It does not tell you the complete sequence of policy decisions that selected one table and one route for a particular flow.

Linux can involve the routing policy database, multiple FIB tables, packet marks, source addresses, input and output interfaces, VRFs, protocol and port selectors, network namespaces and overlay software. A more-specific route in table 52 does not automatically beat a default route in main if policy routing never performs the lookup that would reach table 52.

That distinction is why I do not like starting a routing diagnosis by rebuilding Linux routing logic in Python from a state dump. A simulator can become impressively precise while still being wrong about the behavior of the running machine.

BETTER EVIDENCE

Ask the running kernel the exact routing question.

For a concrete destination, ip route get asks Linux for the resolved path. Supplying the same flow selectors that matter to the incident — source, mark, TOS, interfaces, protocol and ports where relevant — gives you evidence about that flow instead of a generic view of the routing database.

fibmatch adds another useful fact: the FIB route that matched. That lets a diagnostic tool identify the selected prefix without guessing from a dump of every route on the host.

I treat those answers differently from surrounding context. The kernel-selected route is direct runtime evidence. A policy rule whose selectors appear to match is useful context. A route in another table that also contains the destination may be a clue. Those are not the same class of statement.

RPDB

A matching policy rule is not automatically the winning rule.

The routing policy database is processed in priority order, but a selector match does not prove that processing stopped there. A referenced table can fail to return a usable route, a throw route can continue the walk, and suppressors can reject a result.

That means a diagnostic UI should be careful with words such as “matched” and “selected.” If a rule selector can be evaluated from the supplied flow, it is reasonable to say that the selector matches. It is much stronger to say that the rule definitely caused the final route selection.

The same rule applies when information is missing. If a policy rule depends on a selector the operator did not supply, the correct answer may simply be unknown from the available evidence.

CROSS-LAYER

Packet marks make provenance even more important.

A route lookup with --mark 0x42 can tell you what routing decision Linux returns when that mark is present. It does not tell you where the mark came from.

Runtime nftables trace evidence can close part of that gap. If a trace record shows an observed mark or input interface for the packet, a second kernel lookup can use those observed selectors and compare the result with the baseline lookup.

Even then, the conclusion must stay narrow: the re-lookup proves what the kernel returns for the observed selector state. It does not, by itself, prove that Linux rerouted the original packet at that exact nftables hook.

DESIGN RULE

Unknown is better than confidently wrong.

A useful networking diagnostic does not need to pretend it has reconstructed every invisible step. It needs to make provenance visible: what came directly from the kernel, what is derived context, and what still needs investigation.

That is the model behind route-explain: kernel evidence first, correlation second, simulation only where the evidence can support it. The goal is not to replace iproute2; it is to turn several low-level facts into one reviewable answer without changing the machine being diagnosed.

PYTHON CLI / v0.4.1

route-explain

Read-only Linux routing forensics backed by the running kernel.