Centific→
Technical Intern-2 at Centific in Remote
InternshipRemoteRemote$83k–$83k/yr
Skills
pytorchjaxpythoncuda profilingmixed-precision trainingcomputer visionvision-language modelsembodied aiphysical ai3d perceptionexperiment trackingresearch publication
Job Description
Summary: Centific is a frontier AI data foundry that empowers enterprise clients with safe, scalable AI deployment. The Technical Intern will work on building state-of-the-art Vision AI systems, focusing on advancing visual perception, multimodal reasoning, and physical AI, while translating cutting-edge research into production systems.
Responsibilities:
- Advance Visual Perception: Build and fine‑tune models for detection, tracking, segmentation (2D/3D), pose & activity recognition, and scene understanding (incl. 360° and multi‑view)
- Multimodal Reasoning with VLMs: Train/evaluate vision–language models (VLMs) for grounding, dense captioning, temporal QA, and tool‑use; design retrieval‑augmented and agentic loops for perception‑action tasks
- Physical AI & Embodiment: Prototype perception‑in‑the‑loop policies that close the gap from pixels to actions (simulation + real data). Integrate with planners and task graphs for manipulation, navigation, or safety workflows
- Data & Evaluation at Scale: Curate datasets, author high‑signal evaluation protocols/KPIs, and run ablations that make results irreproducible impossible
- Systems & Deployment: Package research into reliable services on a modern stack (Kubernetes, Docker, Ray, FastAPI), with profiling, telemetry, and CI for reproducible science
- Agentic Workflows: Orchestrate multi‑agent pipelines (e.g., LangGraph‑style graphs) that combine perception, reasoning, simulation, and code‑generation to self‑check and self‑correct
Required Qualifications:
- Ph.D. student in CS/EE/Robotics (or related), actively publishing in CV/ML/Robotics (e.g., CVPR/ICCV/ECCV, NeurIPS/ICML/ICLR, CoRL/RSS)
- Strong PyTorch (or JAX) and Python; comfort with CUDA profiling and mixed‑precision training
- Demonstrated research in computer vision and at least one of: VLMs (e.g., LLaVA‑style, video‑language models), embodied/physical AI, 3D perception
- Proven ability to move from paper → code → ablation → result with rigorous experiment tracking
Preferred Qualifications:
- Experience with video models (e.g., TimeSFormer/MViT/VideoMAE), diffusion or 3D GS/NeRF pipelines, or SLAM/scene reconstruction
- Prior work on multimodal grounding (referring expressions, spatial language, affordances) or temporal reasoning
- Familiarity with ROS2, DeepStream/TAO, or edge inference optimizations (TensorRT, ONNX)
- Scalable training: Ray, distributed data loaders, sharded checkpoints
- Strong software craft: testing, linting, profiling, containers, and reproducibility
- Public code artifacts (GitHub) and first‑author publications or strong open‑source impact
Required Skills: PyTorch, JAX, Python, CUDA Profiling, Mixed-Precision Training, Computer Vision, Vision-Language Models, Embodied AI, Physical AI, 3D Perception, Experiment Tracking, Research Publication