Connecting to live activity…

VLM (Vision Language Models) Bootcamp

Multimodal models that connect pixels and language, covering architecture, data, and evaluation.

8 weeks·Hybrid·Cohort format

Overview

Explore vision-language architectures, multimodal datasets, instruction tuning for VLMs, and evaluation beyond single-number scores. The bootcamp is built for engineers who want to work where perception and language meet.

Who it's for

  • CV and NLP engineers crossing into multimodal systems
  • Product engineers building search and assistants with images
  • Candidates targeting multimodal research/engineering roles

What you will learn

Curriculum breakdown, from foundations to portfolio-ready work.

Phase 1

Multimodal foundations

How VLMs represent and fuse modalities.

  • Image encoders + language decoders
  • Alignment objectives and datasets
  • Prompting multimodal models effectively
Phase 2

Training & adaptation

Practical paths to specialize models.

  • Instruction tuning for VL tasks
  • Data curation and toxicity pitfalls
  • Evaluation: VQA, captioning, and grounding basics
Phase 3

Applications

Ship a coherent multimodal feature.

  • Retrieval with images and text
  • Guardrails for multimodal outputs
  • Capstone: multimodal assistant with eval cards

Key features

  • Live paper walkthroughs + coding labs
  • Hands-on multimodal projects
  • Mentorship on experiment design
  • Peer cohort reviews
  • Interview prep on multimodal trade-offs

Projects

Image+text retrieval

Build a small retrieval stack with qualitative failure analysis.

Instruction-tuned VLM task

Adapt behavior for a focused use case with evaluation examples.

Outcomes

Skills gained

End-to-end understanding of multimodal stacks and failure modes.

Job readiness

Articulate architecture and data choices clearly.

Portfolio

Multimodal demos with eval write-ups.

Your mentor

Marcus Chen

Senior Multimodal Research Engineer

11+ years across vision and NLP

Marcus focuses on evaluation rigor and dataset quality, the details that separate demos from products.

VLMsDataEval

Duration & schedule

Total duration

8 weeks

Weekly commitment

11–15 hours

Cohort format

Weekly live sessions; async readings and labs.

Ready to join the next cohort?

Secure your seat or request the full syllabus. We'll confirm prerequisites and start dates.

← All bootcamps