Free and open for anyone to learn or teach · code on GitHub

Module 3 of 8

Working with AI Images, Audio, and Video

One 3-hour session · a short lecture, then a long hands-on lab

Overview

How does a text prompt become a photorealistic image? This session demystifies diffusion models (the technology behind AI image generators), plus speech-to-text, text-to-speech, and voice cloning. We finish with the uncomfortable flip side: deepfakes, and how to spot them.

Learning objectives

  • Describe how diffusion models turn noise into images, in plain language.
  • Generate images with well-crafted prompts and iterate toward a goal.
  • Use speech-to-text to caption a video and text-to-speech to voice one.
  • Identify common signs of AI-generated and deepfaked media.

Session agenda

Materials

  • Free image generation access (e.g. Bing Image Creator, Gemini, or similar).
  • A short video clip to caption (yours, or one provided with the module).
  • Headphones for the audio portions of the lab.

Homework

Generate one image you are proud of and save the full prompt history that got you there. Find one real-world deepfake news story and bring a one-sentence summary.

Detailed slides, lab worksheets, and demos for this module are in progress and will be posted here as they are written.