Autonomous intent-to-fabric operations. You state intent in plain language; a multi-agent tier decomposes it, allocates identifiers and submits declarative resources; Kubernetes controllers reconcile them onto a live SONiC EVPN/VXLAN fabric and keep them that way -- repairing drift, surviving component failure, and releasing what it claimed when intent is withdrawn.
The agent tier is not an add-on. It is how the network is driven: the fabric, the controllers and the agents are three parts of one closed loop, with gNMI telemetry feeding back into it.
Everything below is self-contained: the instructions live here rather than in a separate specification.
Built with Specstride. Not a badge — the method. Every spec, plan, task and verification gate in this lab was driven through Specstride's spec-driven pipeline: the autonomous tier, the controllers and the fabric weren't hand-waved into existence, they were specified into existence, then proven against the running system. A network that writes itself deserves a process that does too — spec it, stride it, ship it.
Full walkthrough (~6½ min, 1.7x) — the intent tier end to end: three services
provisioned from plain-language prompts typed into the operator console (a
vlan, an ip-vrf with its prefix, and a mac-vrf stretched across
both leaves), each one confirmed through the mapper and allocator agents,
reported deployed, and then proven in the terminal with kubectl (the
Network resource Ready, its ApplySucceeded event, its spec) and inside
the SONiC leaf itself (CONFIG_DB, the bridge, the tenant VRF's L3VNI and
Type-5 EVPN route, the VNI tunnel map on both leaves).
Every "deployed" claim in the recording was re-verified from kubectl JSON;
the collected evidence is in
docs/media/agentic-netops-intent-tier-demo-evidence.json.
Two spines, two leaves and four clients in containerlab, wired as a Clos with a dual-stack eBGP underlay and an EVPN/VXLAN overlay.
Live gNMI telemetry from the fabric: gnmic subscribes to each SONiC node's DBs, Prometheus scrapes gnmic directly, Grafana renders it.
The intent tier's console during a real run, wired to the live supervisor:
workers reachable over SLIM, Compass gpt-5 reached through the LiteLLM
gateway, and the mapper's own interpretation of the request shown before
anything is allocated. The scenario cards are the constructs the supervisor
itself advertises on GET /suggested-prompts — vlan, mac-vrf, ip-vrf, acl —
not a separate hard-coded list, and every card names ports this site actually
has. The divider between the workflow canvas and the conversation is draggable
(mouse, touch, or arrow keys); the chosen split is remembered per browser.
The end of the same transaction. The submitted payload is authoritative and
exists only because every apply succeeded; when convergence is still in flight
at the deployer's watch bound the console says exactly that, names the resource,
and points at the status tool — it does not report a success it has not
observed.
What works and what does not: in the console, a provisionable request flows
end to end — classification, interpretation, allocation, your two explicit
confirmations, then a real deployment transaction against the cluster (translate
pod-local, server-side dry-run, deterministic apply, rollback on failure,
convergence watch) and a truthful submitted report — and the southbound
closes the loop: a sonicprovider Network controller renders the accepted
Network onto the SONiC fabric through the host-side fabric-executor and flips
the Network's Ready condition to True only after per-node verification
passes. All four constructs converge. Each of vlan, mac-vrf, ip-vrf and acl
has been driven from plain language to Ready=True on this lab (2026-09-05),
with the VRF/VXLAN/SVI/vtep/access state and the EVPN control plane asserted on
both leaves as applicable. Until 2026-09-04 only the routed construct could ever
converge: L2 services failed at the fabric after the objects were on the
cluster and after the
deployer had reported a successful submission. What was broken in each, and what
is still limited (IPv6 IRB gateways), is recorded in
docs/INTENT_TIER_SERVICE_TYPES.md.
An endpoint naming a node or port the site does not have is refused by the
translator before anything is submitted, with the site's real names listed,
instead of stranding an unrenderable Network on the cluster. And a converged
service is re-applied and re-verified every five minutes, so Ready=True is a
statement about the fabric now rather than about the moment it first converged:
drift — a device manager that died, a reboot, another service's teardown taking
a shared device with it — is repaired, and said so. A single-shot
POST /agent/prompt/stream, however, stops at the supervisor's own iteration
bound, because the graph is built to provision only "after your two explicit
confirmations" and a one-shot request cannot supply them. The refusal is logged
as audit refuse ... reason=request bound — the tier declining
to act, not failing. The transaction contract, including what counts as a
reportable failure at each phase, is specified in
docs/INTENT_TIER_DEPLOYMENT_TRANSACTION.md.
Live gNMI telemetry: gnmic subscribes to each SONiC node's DBs, Prometheus scrapes gnmic directly, Grafana renders it.
| Piece | What it is |
|---|---|
| SONiC fabric | 2 spines, 2 leaves, 4 clients in containerlab; BGP underlay, EVPN/VXLAN overlay |
| Controllers | SRv6Service CRD + provider, built from vendored Go source |
| Observability | OpenTelemetry collector, Prometheus, Grafana with fabric dashboards |
| Intent tier | AGNTCY supervisor + mapper/allocator/deployer agents over A2A/SLIM; the deployer runs the deployment transaction (translate → dry-run → apply → rollback → convergence watch) and reports truthfully |
Read docs/DEPENDENCIES.md first — it lists the host tooling and two traps that will otherwise cost you time:
provision.shdry-runs against your current kubectl context. Point it somewhere harmless or delete stale clusters first, or you get a confusingnamespaces "kubenet-system" not found.kubectl topneeds metrics-server, which kind does not install. Anything that measures resource usage silently returns nothing without it.
For a guided walk-through — including driving the agents from plain language — see TUTORIAL.md.
# bring the fabric up (~30-40 min on first run: image pulls + controller build)
./scripts/provision.sh --profile sonic-vs --cluster-name agentic-netops
# verify
./tests/integration/fabric_verify.sh # BGP sessions, EVPN routes, overlay data path
make verify-pins # every image/binary matches versions.lock.yaml
kubectl --context kind-agentic-netops get pods -A
# tear down (idempotent; safe to re-run)
./scripts/off.sh --delete-kind trueAdd the agent tier:
# LLM provider: copy the example, uncomment one provider block, fill in the key.
# LLM_MODEL is always the model variable; its prefix picks the LiteLLM provider.
# .env is gitignored and becomes Secret/llm-provider at provision time.
cp .env.example .env
$EDITOR .env # e.g. LLM_MODEL=openai/gpt-5, OPENAI_API_KEY=..., OPENAI_BASE_URL=https://api.openai.com/v1
./scripts/provision.sh --profile sonic-vs --cluster-name agentic-netops --with-intent-tier
kubectl --context kind-agentic-netops -n agentic-netops-agents get deploy
# UI on http://localhost:30000These are real, reproduced, and documented rather than hidden:
- The former SONiC ASan/L2-VNI and EVPN Type-5 limitations (D-A2/D-A3) were
resolved on 2026-09-04 by the clean
sonic-vs-gnmi:202505-v1image. The unwaived fabric gate now verifies Type-2/3/5, remote VTEPs, and overlay traffic. - Two pinned images have no local build step (
grafana/flow-plugin,ghcr.io/agentic-netops/topology-generator). Provisioning warns rather than fails; the dependent workload ends inImagePullBackOff. (The intent tier's six images — supervisor, mapper, allocator, deployer, translator, UI — all build locally fromdocker/Dockerfile.*; noteintent::installskips the docker build when the image tag already exists locally unlessINTENT_TIER_REBUILD=true, so source changes need a rebuild or a forced rebuild to reach the cluster.) - IPv6 IRB gateways may not originate their Type-5 route. An IRB carries only
the address families the operator asked for; when IPv6 is requested, this
sonic-vs FRR build sometimes registers the global address as a kernel rather
than a connected route, and
redistribute connectedthen never originates it. The service reportsReady=Falsenaming the missing route rather than claiming success. See docs/INTENT_TIER_SERVICE_TYPES.md. docs/INTENT_TIER_OPS_READINESS.mdcontains resource figures that were never measured. They were produced before a cluster existed and before metrics-server was installed; real values differ by large factors. Re-measure before relying on that document.
scripts/ provision.sh, off.sh, and lib/ (containerlab, rbac, qualify, intent_tier)
lab/ containerlab topology and the sonic-vs profile bootstrap
deploy/ Kubernetes manifests: controllers, kubenet, observability, agents
controllers/ SRv6Service and sonicprovider Network controllers (Go); see README-CONTROLLERS.md
tests/ integration (fabric_verify, cycles_runner) and unit suites
agents/ intent tier: supervisors, provisioning workers, test corpora
docs/ operations, security audit, dependencies, known defects
versions.lock.yaml every image and binary pin; enforced by `make verify-pins`
Jumbo MTU — the lab standardises on underlay MTU 9216. VXLAN effective payload is 9166 (IPv4) and 9162 (IPv6); with 3 SRv6 SIDs it is ~9120. Acceptance tests size packets to avoid fragmentation.
Supply chain — make verify-pins fails if any running image or binary drifts from
versions.lock.yaml.



