top of page

Full-Stack AI Infrastructure for a Startup

  • Case Study
  • Jul 12
  • 3 min read

A deep-tech AI and robotics startup needed a complete infrastructure foundation built from nothing.

Dayo-Tech delivered the full stack, 1.5 PB of parallel storage, a DGX-powered GPU cluster, and a 200 Gb GPUDirect fabric, engineered for high-performance AI.


The Decision


The Customer, a deep-tech AI and robotics startup, needed a complete IT and AI foundation built from nothing, with high-performance GPU compute at its core. A startup at this stage faces a stark choice: assemble infrastructure piecemeal, absorbing the delays and integration gaps that come with it, or bring in a single partner able to design and deliver the entire stack at once.

The requirement was not simply "some infrastructure." It was high-performance GPU compute for demanding AI workloads, with storage, networking, monitoring, and security all engineered to keep that compute fed and protected. Building from a standing start meant every layer had to be designed deliberately and delivered together, which is what drove the decision to engage a full-stack partner rather than piece the environment together over time.


The Architecture


Dayo-Tech engineered the environment around data throughput, because in AI training the path between storage and GPUs is often where performance is won or lost.

At the storage layer sits a 1.5 PB intelligent, high-performance parallel storage system from DDN. The compute layer is a SLURM-scheduled cluster combining NVIDIA DGX B200 systems with ASUS ESC8000A servers equipped with 64 NVIDIA RTX 6000 Ada GPUs. Connecting them is a 200 Gb InfiniBand fabric with GPUDirect, engineered to move data between storage and GPUs with minimal overhead, removing the bottleneck that would otherwise starve high-performance GPUs of data.

Observability and security were built in from the start, not added later. Monitoring came through Zabbix and Grafana; the environment was protected with endpoint detection and response (EDR) and a SOC-as-a-Service capability. Throughout, Dayo-Tech worked hands-on with the Customer's development and IT teams.


The Platform


The result is a complete, high-performance AI environment. SLURM schedules GPU workloads; the 1.5 PB DDN storage feeds them over the 200 Gb InfiniBand fabric via GPUDirect; Zabbix and Grafana provide observability; and EDR together with SOC-as-a-Service protect the environment. A workload moves through the platform scheduled, fed by high-throughput storage, continuously observed, and secured, with no obvious weak link between those stages.

For a fast-moving startup, that completeness is the value. Rather than stitching together compute, storage, and security separately over time, the company received an integrated foundation, delivered and supported as one.


Technology Stack


  • Compute / GPU: NVIDIA DGX B200; ASUS ESC8000A servers; 64 × NVIDIA RTX 6000 Ada GPUs

  • Storage: DDN parallel storage system (1.5 PB)

  • Networking: 200 Gb InfiniBand fabric with GPUDirect

  • Automation / DevOps: SLURM cluster scheduling

  • Security: Endpoint detection and response (EDR); SOC-as-a-Service

  • Monitoring: Zabbix; Grafana


The Impact


The impact is the build itself: a 1.5 PB parallel storage system, a GPU cluster combining DGX B200 systems and ASUS servers with 64 RTX 6000 Ada GPUs, and a 200 Gb GPUDirect fabric engineered for high-throughput AI, delivered from a standing start, with hands-on support throughout. In short order, the Customer went from no infrastructure to a production-grade AI environment operating at real scale.


Closing


This engagement is a clear example of Dayo-Tech acting as a startup's infrastructure partner from the very first server, designing for real performance rather than raw capacity, building in security and observability from day one, and staying close to the team as it scaled. It's the kind of full-stack, high-performance foundation deep-tech companies need when there's no room for infrastructure to become the bottleneck.

bottom of page