MLX, CoreML)PhD with research in efficient multimodal reasoning, LLM reasoning, model compression, efficient inference/decoding, or lightweight VLM architectures Experience training or post-training vision-language models end-to-end - connector/projector design, visual instruction tuning, resolution and token-budget trade-offs, small-model recipes Hands-on experience with reasoning techniques: chain-of-thought distillation and compression, latent/implicit reasoning, reward-guided decoding, RL for reasoning (GRPO, RLVR, STaR), or test-time compute allocation Expertise in decoding and serving optimizations: speculative decoding, structured/grammar-constrained generation, KV-cache quantization and eviction, continuous batching, long-context inference Experience combining LLMs with real-time perception models - 3D reconstruction and geometry (VGGT, DUSt3R/MASt3R-style, SLAM, monocular depth), human pose and body/hand mesh recovery (SMPL-family), detection, segmentation, or tracking - and with spatial or 3D-grounded reasoning and embodied/spatial VQA Experience deploying LLM or multimodal models on mobile or edge hardware (CoreML, MLX, TensorRT-LLM, or equivalent), with attention to ANE/GPU kernel and memory constraints Experience with quantization-aware training, mixed-precision inference, and knowledge distillation for vision-language models Familiarity with efficient vision encoders and self-supervised/joint-embedding pretraining (V-JEPA, I-JEPA, MAE, DINO/DINOv2, SigLIP, CLIP), Mamba/SSM vision backbones, or streaming architectures for real-time video with fixed memory budgets Interest in Video-LLMs, long-video reasoning, and world models for prediction and planning Publication record in top-tier venues is a plus (NeurIPS, ICML, ICLR, CVPR, ECCV, ACL, MLSys, ICRA, etc.). Rather than consuming pixels alone, the VLM should be able to invoke and reason over the outputs of specialist on-device vision models - feed-forward 3D scene and geometry estimators (VGGT-style reconstruction, depth, camera pose), human body and hand mesh/pose recovery, object detectors, localizers, and trackers - and fuse those structured, metric outputs into its reasoning about the scene.