ByteDance→
Research Intern (Video Data Compression and… at ByteDance · San Diego
InternshipOn-siteSan Diego, CA$119k–$119k/yr
Skills
video compressionimage compressiongaussian splatting-based codingtoken compressionlarge multimodal modelspythonpytorchc++deep learning frameworkstensorflowyololarge language modelsvision-language modelsgpu accelerationcudacollaborative mindset
Job Description
Summary: ByteDance is a leading technology company known for its innovative products like TikTok and CapCut. They are seeking a Research Intern to contribute to their Multimedia Lab by designing and optimizing algorithms for video data compression and large multimodal model applications.
Responsibilities:
- Design, develop, and optimize innovative algorithms for data compression, processing, and large multimodal model applications, including but not limited to 2D video, multiview video, point clouds, Gaussian splatting–based or NN-based coding, token compression, and KV cache optimization
- Stay up to date with state-of-the-art techniques through standardization activities or leading conference/journal publications
- Build prototypes and demonstrations, and contribute to technical reports, publications, and patent filings
Required Qualifications:
- Current Ph.D. student in computer science/electrical engineering/mathematics/statistics/data science and related disciplines
- Strong Computer Science fundamentals (algorithms, data structures, software design) and problem-solving skills
- Familiar with token/image/video coding and processing or large multimodal models
- Familiar with Python, PyTorch
- Familiar with C/C++
- Collaborative mindset, with solid written and verbal communication skill
Preferred Qualifications:
- Good understanding of state-of-art compression algorithms
- Rich experience and interest in video/image coding standards (e.g., H.263/264/265/266, MPEG-2/4, JPEG, JPEG 2000, AV1, AV2, AVS1/2/3, etc.)
- Proficient in deep learning frameworks such as PyTorch, TensorFlow, and YOLO
- Hands-on experience with large language models (LLMs) and vision-language models (VLMs), including LoRA, diffusion models, and VQA tasks
- Experience with model training, fine-tuning, and evaluation pipelines
- Familiar with GPU acceleration and CUDA environment configuration
Required Skills: Video compression, Image compression, Gaussian splatting-based coding, Token compression, Large multimodal models, Python, PyTorch, C++, Deep learning frameworks, TensorFlow, YOLO, Large language models, Vision-language models, GPU acceleration, CUDA, Collaborative mindset
Internship Start Date: Start in 2026
Benefits: Interns have day one access to health insurance, life insurance, wellbeing benefits and more., Interns also receive 10 paid holidays per year and paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year)., Interns who are not working 100% remote may also be eligible for housing allowance.
Benefits
Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
Interns also receive 10 paid holidays per year and paid sick time (56 hours if hired in first half of year, 40 if hired in second half of year).
Interns who are not working 100% remote may also be eligible for housing allowance.