Multimodal AIAdvanced Level65 Hours Live

Computer Vision & Multimodal Threat Intelligence

Adversarial patch attacks on vision transformers (ViT), multimodal prompt injection, and synthetic deepfake detection pipelines.

Vision Transformer Adversarial Patches
Cross-Modal Prompt Injection in GPT-4V
Deepfake Forensic Watermarking
65 Hours Practical Workload
1 Core Modules
1 Sandboxed Labs
Cryptographic TS-ID Verifiable

Course Overview & Objectives

Multimodal models accept images, audio, and video alongside text. Discover how adversarial optical illusions trick Vision Transformers, embed hidden prompt instructions within visual textures, and construct deepfake forensic watermarking pipelines.

What You Will Master

  • Generate physical and digital adversarial patches that fool YOLO and Vision Transformers
  • Embed visual prompt injection payloads into images that alter multimodal LLM outputs
  • Implement frequency-domain spectral analysis to detect synthetic AI deepfakes
  • Deploy robust provenance watermarking with C2PA open metadata standards

Prerequisites

  • Python & PyTorch basics
  • Understanding of convolutional or transformer architectures

Platforms & Tools Covered

PyTorchOpenCVHugging Face TransformersYOLOv8Albumentations

Detailed Curriculum Modules

1 modules structured from foundational theory through complex adversarial execution.

65 Total Workload Hours
MODULE 01

Vision Transformer Security & Adversarial Patches

1 Lessons

Attention map disruption, universal adversarial perturbations, and patch generation.

Training an Evasion Patch to Misclassify Object Detectors
55m

Hands-on Virtual Sandbox Labs

Zero local hardware dependencies. Provisioned in cloud containers via browser terminal.

LAB 01~60 mins

Visual Prompt Injection via Steganography

Hide prompt instructions inside image pixel noise that commands GPT-4V to output attacker secrets.

Skills Tested:Multimodal AI, Steganography

Faculty & Lead Instructor

Direct weekly instruction, live office hours, and code-review feedback.

PK

Piya Kohli

Thread Security Education

Multimodal AI Scientist

Leading computer vision research on adversarial image robustness and synthetic media authentication.

Frequently Asked Questions

Everything you need to know about scheduling, cohort admissions, and lab access.

Do we use GPUs for labs?

Yes! Cloud sandbox environments include dedicated GPU instances for neural network training and inference.

Ready to Master Computer Vision & Multimodal Threat Intelligence?

Join the upcoming cohort. Seats are limited to maintain a high faculty-to-student ratio and rigorous sandbox feedback.