Experience building Retrieval-Augmented Generation (RAG) pipelines and working with vector databases (pgvector, Qdrant, Weaviate). Deep understanding of self-hosted AI infrastructure, including model formats, quantization, GPU memory management, batching, and inference optimization.