Cisco Architecture Proof of Concept

The Edge
Intelligence
Grid.

We transform Cisco Catalyst Switches and CW917X Access Points into a decentralized AI execution farm via Pipeline Parallelism. Zero dedicated GPUs. Total air-gapped sovereignty.

CATALYST L0 L1 L14 L15

Beating the 100MB IOx limit.

Cisco CW917X APs have powerful ARM64 CPUs, but `cgroups` restrict memory to roughly 100MB. A 1B parameter model needs 2.2GB. Here is the math behind the magic.

01. Quantization

INT4 GGUF

We compress Llama-3.2-1B weights to 4 bits (Q4_K_M). Total model footprint plummets from 2.2GB to ~800MB with near-zero perplexity loss.

02. Pipeline Parallelism

Layer Slicing

The Llama-3.2-1B model contains 16 Transformer Layers. Instead of loading it fully, we slice the 800MB model horizontally across the network.

Total Size 800 MB
Total APs รท 16 Layers
Memory per Node ~50 MB
root@cat9300x-orchestrator:~

$ divaid-cli start --model llama3.2-1b --quant INT4

[INFO] Initializing Swarm Orchestrator...

[INFO] Discovering IOx Nodes on VLAN 100...

[OK] Found 16 active CW9178 Access Points.

[INFO] Sharding Llama-3.2-1B.gguf (812MB) into 16 parts...

[OK] Layer 0 deployed to 10.0.0.11 (48.1 MB)

[OK] Layer 1 deployed to 10.0.0.12 (48.1 MB)

... [Layers 2-14 deployed] ...

[OK] Layer 15 deployed to 10.0.0.26 (48.1 MB)

> Swarm Ready. Awaiting Prompts.

03. Execution

Success

50MB safely clears the Cisco 100MB hard limit. The container boots, the weights load into RAM, and the AP becomes a permanent Layer executor for the swarm.

Like a Blockchain for AI Operations.

Traditional inference requires a monolithic GPU holding the entire model in VRAM. We shatter this assumption.

Think of blockchain: millions of miners process tiny chunks of cryptographic work that chain together. Divaid uses your ceiling APs exactly like this.

Each AP computes its dedicated transformer layer and passes a microscopic 4KB hidden-state tensor to the next AP via Cisco's mGig backplane in microseconds.

Execution Pipeline
1. Orchestrator Tokenizer & Embeddings
2. AP Node 01 Attention + FFN (L0)
gRPC Tensor 4KB
3. AP Node 02 Attention + FFN (L1)
...
N. Output LM Head Token