Overview
Explore vision-language architectures, multimodal datasets, instruction tuning for VLMs, and evaluation beyond single-number scores. The bootcamp is built for engineers who want to work where perception and language meet.
Who it's for
- CV and NLP engineers crossing into multimodal systems
- Product engineers building search and assistants with images
- Candidates targeting multimodal research/engineering roles
What you will learn
Curriculum breakdown, from foundations to portfolio-ready work.
Multimodal foundations
How VLMs represent and fuse modalities.
- Image encoders + language decoders
- Alignment objectives and datasets
- Prompting multimodal models effectively
Training & adaptation
Practical paths to specialize models.
- Instruction tuning for VL tasks
- Data curation and toxicity pitfalls
- Evaluation: VQA, captioning, and grounding basics
Applications
Ship a coherent multimodal feature.
- Retrieval with images and text
- Guardrails for multimodal outputs
- Capstone: multimodal assistant with eval cards
Key features
- Live paper walkthroughs + coding labs
- Hands-on multimodal projects
- Mentorship on experiment design
- Peer cohort reviews
- Interview prep on multimodal trade-offs
Projects
Image+text retrieval
Build a small retrieval stack with qualitative failure analysis.
Instruction-tuned VLM task
Adapt behavior for a focused use case with evaluation examples.
Outcomes
Skills gained
End-to-end understanding of multimodal stacks and failure modes.
Job readiness
Articulate architecture and data choices clearly.
Portfolio
Multimodal demos with eval write-ups.
Your mentor
Marcus Chen
Senior Multimodal Research Engineer
11+ years across vision and NLP
Marcus focuses on evaluation rigor and dataset quality, the details that separate demos from products.
Duration & schedule
8 weeks
11–15 hours
Weekly live sessions; async readings and labs.
Ready to join the next cohort?
Secure your seat or request the full syllabus. We'll confirm prerequisites and start dates.
← All bootcamps