Overview
This is an incredible opportunity to be part of a company that has been at the forefront of AI and high-performance data storage innovation for over two decades. DataDirect Networks (DDN) is a global market leader renowned for powering many of the world''s most demanding AI data centers, in industries ranging from life sciences and healthcare to financial services, autonomous cars, Government, academia, research and manufacturing.
"DDN''s A3I solutions are transforming the landscape of AI infrastructure." - IDC
"The real differentiator is DDN. I never hesitate to recommend DDN. DDN is the de facto name for AI Storage in high performance environments" - Marc Hamilton, VP, Solutions Architecture & Engineering | NVIDIA
DDN is the global leader in AI and multi-cloud data management at scale. Our cutting-edge data intelligence platform is designed to accelerate AI workloads, enabling organizations to extract maximum value from their data. With a proven track record of performance, reliability, and scalability, DDN empowers businesses to tackle the most challenging AI and data-intensive workloads with confidence.
Our success is driven by our unwavering commitment to innovation, customer-centricity, and a team of passionate professionals who bring their expertise and dedication to every project. This is a chance to make a significant impact at a company that is shaping the future of AI and data management.
Our commitment to innovation, customer success, and market leadership makes this an exciting and rewarding role for a driven professional looking to make a lasting impact in the world of AI and data storage.
Job Description
DDN is expanding our Enterprise and Sovereign AI Solution offerings, for example Hyperpod - a turnkey NVIDIA AI Data Platform built on DDN Infinia storage, NVIDIA AI Enterprise (NVAIE), and Supermicro reference hardware, optimized for inference and RAG workloads. Our support organization is deep on storage (Infinia, EXAScaler); we are now hiring an AI platform specialist to lead supportability and enablement for the AI side of the stack - NVIDIA AI Enterprise services (NIMs, NeMo, Triton, GPU Operator, licensing), vector databases (initially Milvus), RAG/agentic workflows, and the high‑performance storage and networking fabric that underpins them.
You will be a trusted technical advisor within Support and across OEM and NVIDIA partner teams, combining the mindset of a solutions architect (architecture, reference patterns, PoCs, reusable assets) with that of a L3 support engineer. You'll help DDN and our partners operate AI Data solutions as a cohesive AI platform, not just a collection of components.
Key Responsibilities
Platform support
Runbooks, diagnostics, and supportability
Enablement, PoCs, and reusable assets
Design feedback, readiness, and cross‑functional leadership
Required Experience & Skills
Technical
5+ years in Linux‑based infrastructure roles (SRE, MLOps, platform engineering, or L2/L3 support) supporting production systems; 8+ years total technical experience preferred.
Strong hands‑on experience with containers and Kubernetes (Docker/containerd, Helm, Operators; debugging pods, DaemonSets, CSI, CNI, and ingress/load balancers).
Demonstrated experience operating GPU‑accelerated workloads in production:
NVIDIA GPUs, drivers, CUDA concepts, GPU utilization/perf triage
NVIDIA GPU Operator and Kubernetes‑based GPU lifecycle management
Familiarity with DGX / HGX or similar GPU cluster platforms.
Practical experience with AI storage and networking for HPC/AI clusters:
High‑performance storage systems (e.g., EXAScaler/Lustre, GPFS, Ceph, distributed object storage, enterprise NAS/SAN).
RDMA‑accelerated and/or high‑speed Ethernet/InfiniBand networking, including fabrics, switch topologies, and large‑scale deployments.
Hybrid cloud or cloud‑adjacent patterns (Kubernetes CSI, cloud‑native fabrics, data locality).
Experience with one or more vector databases (Milvus, Qdrant, Pinecone, pgVector, OpenSearch/Elasticsearch vectors, etc.), including schema design, ingestion, and operations.
Solid understanding of RAG and Generative AI workflows: embeddings, retrieval, reranking, prompt design, context management, and how these interplay with vector search and GPU inference at scale.
Familiarity with NVIDIA AI Enterprise components and toolchain, for example:
NVIDIA NIM inference microservices
NVIDIA NeMo framework / NeMo Retriever / NeMo Curator
Triton Inference Server, TensorRT / TensorRT‑LLM, CUDA libraries
NVIDIA blueprints for enterprise RAG and agentic AI.
Experience designing, operating, or supporting MLOps / GenAI pipelines: CI/CD for models, deployment strategies, canarying/rollback, GPU resource management, monitoring and alerting for AI services.
Strong diagnostic skills across Linux, containers, Kubernetes, GPUs, storage, and networking; able to quickly narrow fault domains and propose experiments or configuration changes.
Support, architecture, and stakeholder skills
Preferred Qualifications
What Success Looks Like in This Role
Within 6-12 months, a successful AI Data Platform Solutions Architect will have:
Salary Range for this role: $175,000 - $200,000
DDN
DDN has a very strong orientation towards these 4 characteristics and any successful employee will demonstrate these capabilities:
Self-Starter - Takes independent action to identify and solve problems. Seeks out relevant information needed to make decisions. Gets involved with new initiatives.
Success/Achievement Orientation - Delivers quality results consistently. Targets, achieves (or exceeds) measurable results. Sets challenging goals, focuses on critical priorities, and is accountable.
Problem Solving - Recognizes problems and responds with a systematic assessment that identifies and addresses cause of issue. Practical, realistic, and resourceful.
Innovative - Builds and improves key business processes that enhance the effectiveness of DDN. Generates new ideas, challenges the status quo, and solves problems creatively.
DataDirect Networks, Inc. is an Equal Opportunity/Affirmative Action employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity, gender expression, transgender, sex stereotyping, sexual orientation, national origin, disability, protected Veteran Status, or any other characteristic protected by applicable federal, state, or local law.
#LI-Remote