Phase 3: Giving an AI Read-Only Access to the Fabric via MCP
Building two MCP servers - one auto-generated from NetBox's OpenAPI schema, one hand-built on Netmiko - and testing whether an AI actually notices something real on the fabric, or just reports 0% packet loss and calls it done.
Phase 2 ended with a question: once an AI can run show commands against this fabric itself, does it actually notice something real - the TTL difference between an L2VNI-bridged ping and an L3VNI-routed one - or does it just report 0% packet loss and call the job done?
To even ask that question, the AI needs a way to reach the fabric in the first place. That’s this phase: two MCP (Model Context Protocol) servers - one giving an AI read access to NetBox as inventory, one giving it the ability to run read-only commands against the actual leafs and spines over Netmiko - plus everything that broke while building them, because most of what’s worth writing about here is exactly that.
Two servers, two different build styles
NetBox server - generated automatically from NetBox’s own OpenAPI schema:
mcp = FastMCP.from_openapi(
openapi_spec=openapi_spec,
client=client,
name="NetBox MCP",
route_maps=[
RouteMap(methods=["DELETE"], mcp_type=MCPType.EXCLUDE),
],
)
One function call turns NetBox’s entire REST API surface into ~500 callable tools - every object type, every filter. DELETE is excluded at the route level, so there’s no delete-shaped tool anywhere in the generated set, regardless of what permissions the token has.
Starting point adapted from PacketCoders’ walkthrough on dynamically creating MCP servers with FastMCP - everything from “what broke” onward in this post is what happened once I pointed that pattern at a real NetBox instance and a real AI client.
Once it was actually working, asking for devices by site went straight through NetBox’s real API - no code written for this specific query, just the ~500 auto-generated tools already covering it:
Netmiko server - hand-written, because there’s no OpenAPI schema for a CLI. It looks devices up in NetBox via pynetbox, then opens the actual SSH session:
def connect_and_run(device_name, command):
...
connection = ConnectHandler(**connection_data)
output = connection.send_command(command)
connection.disconnect()
return output
Two very different amounts of code to get to ‘a tool the AI can call’ - one function call vs. a hand-rolled server. And they failed in very different ways.
What broke building the NetBox server
Two bugs before it would even start.
A variable that never got reassigned. The schema-fetch path built the cache file correctly but never updated the variable actually passed to from_openapi():
# before
openapi_spec = httpx.get(...) # this is a Response object
openapi_spec.raise_for_status()
SCHEMA_CACHE.write_text(json.dumps(openapi_spec.json())) # cache written correctly...
# ...but openapi_spec itself is still the Response object here
from_openapi() needs a dict, not a Response. The cache file was fine; the variable feeding the actual server wasn’t. Fix: assign the parsed JSON to a variable and use that everywhere.
A schema NetBox emits that JSON Schema’s newer draft rejects. One field (position, a rack-unit value) came back with "exclusiveMaximum": true - valid OpenAPI 3.0 syntax, but not valid under the JSON Schema draft FastMCP validates tool outputs against, which expects exclusiveMaximum to be a number, not a boolean. Every call that touched that field failed schema validation before the data even reached the AI. from_openapi(..., validate_output=False) was the fix - it turns off that particular the output-shape check without affecting anything about how the actual API calls work.
Locking down what the AI can actually see
The first working token was a personal admin account’s - which meant a read-only-looking test could, in principle, see or touch anything in NetBox. Fix: a dedicated mcp-agent NetBox user, non-superuser, granted exactly one permission - view on dcim.device, nothing else.
That’s where I found something I wasn’t expecting: NetBox’s device serializer embeds config_context directly on the device object, and on this platform that includes live credential material - enable-secret hashes, a RADIUS shared key - inline in the same response as the device’s name and IP. Scoping the permission tighter doesn’t hide it, because it’s not a separate object type; it’s a field on the object you’re already allowed to view. The actual fix was at the query level - fields=id,name,platform,primary_ip on every device-list call, so config_context never leaves NetBox in the first place. Least-privilege on the account and field-scoping on every call turned out to be two separate, both-necessary controls, not one.
The Netmiko server’s real bug: two tools quietly sharing state
Two tools, run_netbox_command (inventory) and execute_command (run a command), both read from the same module-level dict:
devices = {} # populated only by run_netbox_command
def connect_and_run(device_name, command):
...
elif device_name in devices: # reads it directly, never populates it
Call execute_command cold, before run_netbox_command has ever run in that process, and every device lookup fails - not with an error that says “inventory not loaded,” but with Device 'L1' not found, which looks exactly like the device doesn’t exist. I spent a few exchanges chasing a token-rotation theory before actually checking this. Fix was one line: if not devices: get_netbox_device() at the top of connect_and_run, so the tool populates its own inventory instead of depending on call order the AI has no way of knowing about.
A safety filter, and the gap in it
execute_command originally ran the same three fixed commands no matter what was asked. Making it accept an arbitrary command meant it needed an actual filter - not just “starts with show”:
BLOCKED_SHOW_PATTERNS = [
"show running-config",
"show startup-config",
"show tech",
]
if normalized.startswith(blocked):
return False
This blocks show running-config correctly. It does not block show run - a completely standard IOS abbreviation that the device itself expands to the same command. And this wasn’t a hypothetical attack: asked in plain language to “show me the running config,” the AI reached for show run on its own, zero adversarial intent required - it’s just idiomatic Cisco shorthand. A blocklist that only recognizes one exact spelling doesn’t survive contact with a device that accepts abbreviations. Fix - check the prefix relationship in both directions, since any valid IOS abbreviation is by definition a text-prefix of the full command:
if normalized.startswith(blocked) or blocked.startswith(normalized):
return False
show run, sh run, show ru - all caught now, without having to enumerate every possible abbreviation by hand.
The filter actually rejecting a request, and show version going through cleanly right after it:
So - does it notice?
Before the VNI check, a simpler live query - PIM neighbors and BGP EVPN state on L1 and L2, just to confirm the AI could actually drive real show commands against real devices, not just NetBox:
With both servers working, I asked it to pull L2VNI state (show nve vni) across all four leafs - the same VNIs this series has been building since Phase 2.
L1, L2, and L4 came back clean. L3’s output for VNI 10100 didn’t:
Interface VNI Multicast-group VNI state Mode VLAN cfg vrf
nve1 10100 239.0.0.100 BD Down/Re L2CP N/A CLI N/A
Every other row - on every leaf - has a real VLAN number and a real VRF in those last two columns. This one has N/A in both, and a state that isn’t Up. Nothing prompted me to go looking for this specifically; it showed up in an otherwise routine inventory pull, and the shape of the row (VNI present, but no VLAN, no VRF, state down) was enough to work out what it meant without needing show running-config to confirm it: L3 still has an NVE member vni 10100 entry advertised into the fabric, but the VLAN it’s supposed to bind to isn’t configured there anymore - a leftover from earlier in the build, not something intentionally removed on purpose.
So it did answer the Phase 2 question, just not the way I expected. I was looking for the TTL difference. Instead, it found a stale config artifact sitting in plain sight: three leafs reported the VNI as Up, and the fourth had Down/Re with N/A for both VLAN and VRF.
The full L2VNI check, plus the MAC-IP mapping that goes with it:
What’s next
The MCP layer is intentionally read-only right now - no way to clear that stale member vni 10100 line, or fix anything else, without deliberately building and gating a write path. That’s an open question for a future phase, not an oversight: config-write access is a different risk category entirely, and I want a real answer for “what does confirmation/approval look like before an AI pushes a change” before building it, not just a bigger blocklist.
Asked directly, it confirms there’s no path to do this at all right now - not a blocked command, just no tool that can:
Code
net-auto-labs/mcp_netbox_netmiko
Resources
- PacketCoders: How to Dynamically Create MCP Servers with FastMCP - the starting point for the NetBox server in this lab
- Model Context Protocol
- FastMCP
- Netmiko documentation
- NetBox REST API
