Cisco Infrastructure Native

The Edge
Intelligence
Grid.

We transform Cisco Catalyst Switches and CW917X Access Points into a sovereign, decentralized AI execution farm. Leveraging asymmetric pipeline parallelism to bypass hardware constraints and execute high-density AI workloads with zero dedicated GPUs.

CATALYST L[n] L[n+1] L[n+k] L[last]
Zero-Trust Execution

Absolute Data Sovereignty.

Consider a financial institution processing highly sensitive PII (Personally Identifiable Information). Cloud-based LLMs leak context via external APIs. Divaid keeps the prompt entirely localized. The inference loop is physically sandboxed within the walls of the branch.

EXTERNAL CLOUD / INTERNET (BLOCKED) AIR-GAP FIREWALL LOCAL BRANCH VLAN (SECURE) CORE AP-VAULT AP-TELLER AP-LOBBY

Every token generated in this simulation travels purely through the physical copper and fiber inside the building.

When a loan officer queries a high-net-worth client's profile, the Orchestrator Core Switch tokenizes the prompt. The subsequent mathematical execution bounces sequentially between AP-Vault, AP-Teller, and AP-Lobby.

The Firewall strictly drops any external inference API calls. The internal VLAN acts as an impenetrable execution substrate.

Beating the IOx memory bottleneck.

Cisco edge infrastructure, such as the CW917X Access Points, possesses robust ARM64 compute capacity. However, Cisco Application Hosting (IOx) restricts active processes via strict cgroups, limiting available RAM to an incredibly tight footprint. Running high-density transformer models natively in this sandbox was deemed impossible. Until now.

01. Deep Quantization

GGUF Encoding

We subject state-of-the-art dense transformer architectures to aggressive INT4 quantization. By dropping precision on the multi-dimensional weight matrices from FP16 to a tightly packed GGUF format, the overall model payload is algorithmically compressed by nearly 65%, with mathematically negligible perplexity degradation.

02. Asymmetric Sharding

Pipeline Parallelism

Even aggressively quantized, the model remains too large for a single IOx container. Our orchestrator dissects the network structurally. Rather than passing parameter weights, we shard the underlying Transformer Block Layers themselves. Each AP node loads only a precise, highly-specific slice of the calculation graph.

Container Payload Sub-block matrix subset
Swarm Distribution N / Active APs
Memory per IOx Node < Cisco cgroup threshold
root@cat9300x-orchestrator:~

$ divaid-cli start --arch Transformer --quant INT4

[INFO] Initiating gRPC Swarm Orchestrator process...

[INFO] Discovering accessible IOx instances via LLDP / VLAN...

[OK] Discovered high-density network edge: AP Cluster active.

[INFO] Generating asymmetrical tensor graph sharding...

[OK] Subgraph block injected into IOx container 10.0.0.11

[OK] Subgraph block injected into IOx container 10.0.0.12

... [Sequential network fabric allocation] ...

[OK] Final LM-head mapped to Orchestrator Core.

> Pipeline established. Listening for inference requests.

03. Edge Execution

Success

The sharded footprint comfortably sidesteps the aggressive OOM-killer logic of the host OS. The container initializes instantly, mapping weights into ARM RAM without relying on swap, turning the AP into a permanent Layer Executor for the swarm.

A Blockchain-inspired Execution Paradigm.

The standard industry approach forces monolithic GPU clusters to handle generation linearly. By treating inference as a series of distributed discrete matrix multiplications, we pivot to a fundamentally different paradigm.

Imagine a cryptographic blockchain: the network avoids centralized computing bottlenecks by forcing millions of disparate nodes to process tiny, fractional hashing tasks that ultimately link together to form a validated chain.

Divaid applies this exact topology to Edge AI. The Cisco access point network acts as the miner pool. A token's hidden-state tensor is serialized and passed across the mGig backplane from node to node, resolving attention layers sequentially without ever leaving your sovereign air-gapped VLAN.

Inference Transport Topology
1. Orchestrator Tokenization & High-D Embeddings
2. Node Cluster Alpha Attention + FFN Computation
gRPC Serialization Payload (~4KB)
3. Node Cluster Beta Sequential FFN Execution
...
N. Resolution LM Head Decode & Yield