LLM Engineer

Talent Software Services, Inc.

  • Cincinnati, OH
  • Today
  • Remote
  • $60.60–$62.87 Per Hour

Highlights

Core Technologies: LLM, SLM, RAG, Generative AI, Agentic AI, Python, Kubernetes, MLflow, Hugging Face, PyTorch, vLLM, KServe, Docker, Vector Databases. Design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities.

Numbers & Facts

LocationCincinnati, OH (
Remote
)
Salary$60.60–$62.87 Per Hour

Description

Job Details

  • Job Title: LLM Engineer
  • Location: Cincinnati, OH
  • Work Location: Remote – USA
  • Duration: 1 year
  • Experience Required: 8+ years
  • Role Category: AI and Automation
  • Education: Any degree
  • Start Date: 15-September-2026

Role Overview

  • Design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities.
  • Build secure, reliable, reusable, and enterprise-ready AI capabilities.
  • Support:
    • Agentic AI workflows
    • AI for SDLC
    • Knowledge retrieval
    • Model evaluation
    • Private AI hosting
    • AgentOps
  • Work closely with:
    • Principal AI Architect
    • AI Engineering Lead
    • Platform Engineers
    • Security teams
    • Enterprise Architecture
    • Product Owners
    • Domain teams

Must-Have Technical Skills

LLM & Generative AI

  • Large Language Models (LLMs)
  • Small Language Models (SLMs)
  • Prompt Engineering
  • Context Engineering
  • Retrieval-Augmented Generation (RAG)
  • Embeddings
  • Semantic Search
  • Agentic AI Patterns
  • Multi-Agent Workflows
  • Tool Calling
  • Function Calling
  • Model Evaluation
  • LLM Observability

Model Engineering

  • Fine-Tuning
  • Supervised Fine-Tuning
  • LoRA
  • QLoRA
  • Quantization
  • Distillation
  • Model Compression
  • Synthetic Data Generation
  • Model Benchmarking
  • Model Selection
  • Model Routing

Model Hosting & Serving

  • Private LLM Hosting
  • On-Prem Model Deployment
  • GPU-Based Inference
  • Model Serving APIs
  • High-Availability Inference
  • Autoscaling
  • Load Balancing
  • Caching
  • Batch and Real-Time Inference

AI Infrastructure & Frameworks

  • Kubernetes
  • Docker
  • Kubeflow
  • KServe
  • Ray Serve
  • MLflow
  • Hugging Face
  • Transformers
  • PyTorch
  • PEFT
  • DeepSpeed
  • NVIDIA NIM
  • Triton Inference Server
  • TensorRT-LLM
  • vLLM
  • TGI
  • SGLang

Programming & Engineering

  • Python
  • TypeScript or JavaScript
  • REST APIs
  • Microservices
  • CI/CD
  • GitHub or Azure DevOps
  • API Design
  • Distributed Systems
  • Cloud-Native Engineering
  • Test Automation

Data & Knowledge Systems

  • Vector Databases
  • Knowledge Graphs
  • Document Processing
  • Metadata Management
  • Data Pipelines
  • Object Storage
  • Enterprise Search
  • Structured and Unstructured Data Integration

Roles & Responsibilities

LLM Application Engineering

  • Build enterprise-grade LLM-powered applications and intelligent agent capabilities.
  • Design reusable LLM patterns, services, APIs, and accelerators.
  • Develop model interaction patterns for:
    • Reasoning
    • Summarization
    • Classification
    • Extraction
    • Planning
    • Decision support
  • Build reusable prompt, context, retrieval, memory, and evaluation components.
  • Support AI-for-SDLC agents across:
    • Requirements
    • Design
    • Coding
    • Testing
    • Security Review
    • Deployment
    • Operations
  • Convert AI use cases into scalable production solutions.

Agent Factory Intelligence Layer

  • Build core intelligence services for enterprise agents.
  • Develop reusable capabilities for:
    • Planning
    • Task decomposition
    • Reasoning
    • Tool usage
    • Agent collaboration
  • Enable agent-to-agent interaction and multi-agent orchestration.
  • Integrate LLMs with:
    • Agent runtimes
    • Tool registries
    • Workflow engines
    • MCP-based gateways
  • Support human-in-the-loop, approval, escalation, and feedback workflows.
  • Improve agent quality, accuracy, safety, and task completion.

Prompt & Context Engineering

  • Design reusable prompt engineering standards, templates, and libraries.
  • Create:
    • System prompts
    • Task prompts
    • Role prompts
    • Guardrail prompts
    • Evaluation prompts
  • Develop context engineering strategies for better grounding, relevance, and personalization.
  • Optimize:
    • Token usage
    • Context windows
    • Memory injection
    • Retrieval inputs
  • Establish prompt versioning, testing, and governance practices.

Retrieval-Augmented Generation (RAG)

  • Design and implement enterprise RAG architectures.
  • Build retrieval pipelines using:
    • Enterprise documents
    • Knowledge repositories
    • Structured data
    • Metadata
  • Optimize:
    • Chunking
    • Embeddings
    • Indexing
    • Ranking
    • Reranking
    • Retrieval strategies
  • Improve grounding, citation quality, precision, recall, and factual accuracy.
  • Build reusable retrieval services for agents and business domains.
  • Partner with data and knowledge management teams to onboard trusted data sources.

LLM / SLM Model Engineering

  • Evaluate, build, fine-tune, deploy, and optimize LLMs and SLMs.
  • Support domain-specific model development using approved datasets.
  • Build supervised fine-tuning and model adaptation pipelines.
  • Apply:
    • LoRA
    • QLoRA
    • Distillation
    • Quantization
    • Model compression
  • Evaluate commercial, open-source, and internally hosted models.
  • Select models based on:
    • Accuracy
    • Latency
    • Cost
    • Data residency
    • Security
    • Operational requirements

Private AI & On-Prem Hosting

  • Build and support private AI capabilities for LLM/SLM hosting.
  • Deploy models across:
    • On-premises
    • Hybrid
    • Private cloud environments
  • Support GPU-enabled model hosting.
  • Optimize latency, throughput, concurrency, resiliency, and GPU utilization.
  • Build secure inference endpoints for internal applications and agents.
  • Support air-gapped and restricted AI environments.
  • Partner with infrastructure and platform teams on private AI hosting.

Model Serving & Inference Optimization

  • Implement scalable model serving using modern inference frameworks.
  • Build high-availability inference architectures.
  • Optimize:
    • Token throughput
    • Response latency
    • Cost efficiency
    • Inference performance
  • Implement:
    • Model routing
    • Load balancing
    • Caching
    • Fallback strategies
  • Support batch and real-time inference.
  • Develop reusable deployment templates for different model families.

LLMOps / ModelOps / AgentOps

  • Build operational practices for managing models and agents throughout their lifecycle.
  • Implement observability for:
    • Prompts
    • Retrieval
    • Model responses
    • Latency
    • Cost
    • Failures
  • Develop evaluation pipelines for regression testing and continuous quality improvement.
  • Monitor:
    • Model drift
    • Response quality
    • Hallucination indicators
    • Safety risks
  • Support CI/CD and release management for:
    • Prompts
    • Models
    • Agents
    • Retrieval pipelines
  • Build dashboards and metrics for AI quality, reliability, adoption, and operational readiness.

AI Evaluation & Benchmarking

  • Define and implement LLM evaluation frameworks.
  • Measure:
    • Accuracy
    • Groundedness
    • Relevance
    • Hallucination rate
    • Toxicity risk
    • Safety compliance
    • Task completion
    • User satisfaction
  • Build automated test suites for prompts, agents, tools, and RAG pipelines.
  • Benchmark models across enterprise use cases.
  • Compare cloud, open-source, and on-prem models based on performance, cost, quality, and risk.
  • Establish quality gates for production AI releases.

Responsible AI, Security & Governance

  • Implement Responsible AI controls in LLM applications and agent workflows.
  • Develop guardrails for:
    • Safe output
    • Tool usage
    • Data access
    • Enterprise policy compliance
  • Support:
    • Model risk management
    • Auditability
    • Transparency
    • Traceability
  • Ensure sensitive data is handled according to security and privacy requirements.
  • Collaborate with Security, Enterprise Architecture, Risk, and Compliance teams.
  • Support model and agent approval and production-readiness governance.

Preferred / Additional Experience

  • Experience deploying open-source models such as:
    • Llama
    • Mistral
    • Mixtral
    • Phi
    • Gemma
    • Qwen
    • DeepSeek
    • Granite
    • Falcon
    • Domain-specific models
  • Experience with GPU infrastructure such as:
    • NVIDIA H100
    • H200
    • B200
    • B300
    • A100
    • L40S
    • GH200
    • AMD MI300X
  • Experience with:
    • Private AI
    • Hybrid AI
    • Air-gapped AI environments
  • Experience in regulated industries such as:
    • Healthcare
    • Financial Services
    • Insurance
  • Experience building:
    • Enterprise copilots
    • AI assistants
    • Agent platforms
  • Experience with:
    • MCP
    • Tool registries
    • Agent runtimes
    • Enterprise integration patterns
  • Experience with Responsible AI, model governance, model risk management, and AI compliance.
  • Experience optimizing AI workloads for:
    • Cost
    • Performance
    • Latency
    • Security

Role / Skills Summary

  • Role: LLM Engineer
  • Essential Skill: LLM
  • Primary Skill: AI and Automation
  • Experience: 8–10+ years
  • Work Location: Remote USA
  • Duration: 1 year
  • Core Technologies: LLM, SLM, RAG, Generative AI, Agentic AI, Python, Kubernetes, MLflow, Hugging Face, PyTorch, vLLM, KServe, Docker, Vector Databases

Similar Jobs

See more jobs