AI Assisted deployment
Intelligence at the Edge with CAPE’s Cognitive Software
The Vision
Europe’s Edge-Cloud landscape is transforming rapidly. Real-time applications demand low latency and high privacy. CAPE meets this shift with a new architecture. We unify cloud automation with edge performance.
Intelligent Orchestration
Modern workloads are not static. They move between cloud and edge dynamically. CAPE’s engine coordinates these complex processes. Tasks are assigned based on cost and latency. We also track real-time resource availability.
The “Cognition” Layer
We introduce Cognitive Infrastructure-from-Code (IfC). Operators describe their needs in natural language. No more manual, low-level configuration. CAPE’s AI interprets the intent automatically. It generates the necessary infrastructure logic. This brings true autonomy to the edge. The CAPE project defines two layers of cognitive abstraction designed to create a self-optimizing edge-cloud platform:
Level 1 – Cognitive Infrastructure from Code (IfC): Focuses on user interaction. It uses adapted Large Language Models (LLMs) to translate natural language descriptions (e.g., “Deploy a database…”) into executable infrastructure code, making the platform accessible to non-experts
Level 2 – Cognitive Application Placement: Focuses on autonomous orchestration. Intelligent agents use real-time telemetry and infrastructure data to decide where workloads should run based on specific policies like minimising latency or energy consumption
This framework uses a hybrid AI strategy, deploying Small Language Models (SLMs) at the edge for real-time decisions and LLMs for complex reasoning.
Hardware-Aware Management
Deploying workloads efficiently requires software that understands the underlying hardware. CAPE’s management layer exposes the full heterogeneity of the platform — CPUs, GPUs, FPGAs and RISC-V accelerators — through a unified, easy-to-integrate interface powered by openMPMC. Rather than treating all compute as equivalent, the Ryax platform continuously profiles available resources and automatically offloads each workload to the architecture best suited for it: latency-critical inference goes to the GPU, energy-constrained background tasks to RISC-V, and data-intensive pipelines to CXL-disaggregated memory pools. The result is an Edge-Cloud platform that adapts to the workload, not the other way around