Overview
How does a text prompt become a photorealistic image? This session demystifies diffusion models (the technology behind AI image generators), plus speech-to-text, text-to-speech, and voice cloning. We finish with the uncomfortable flip side: deepfakes, and how to spot them.
Learning objectives
- Describe how diffusion models turn noise into images, in plain language.
- Generate images with well-crafted prompts and iterate toward a goal.
- Use speech-to-text to caption a video and text-to-speech to voice one.
- Identify common signs of AI-generated and deepfaked media.
Session agenda
- 0:00 – 0:10Warm-up
Study-guide homework debrief: where did the AI shine and where did it fail?
- 0:10 – 1:15Lecture: Diffusion models, speech, and deepfakes
Noise-to-image intuition for diffusion; speech-to-text and text-to-speech; a live voice-cloning demo and the deepfake problem.
- 1:15 – 1:25Break
- 1:25 – 3:00Lab: Create and caption
Generate images from prompts and refine them; auto-caption a short video; experiment with text-to-speech voices.
- 3:00 – 3:10Wrap-up
Gallery walk of generated images; preview AI and data.
Materials
- Free image generation access (e.g. Bing Image Creator, Gemini, or similar).
- A short video clip to caption (yours, or one provided with the module).
- Headphones for the audio portions of the lab.
Homework
Generate one image you are proud of and save the full prompt history that got you there. Find one real-world deepfake news story and bring a one-sentence summary.
Detailed slides, lab worksheets, and demos for this module are in progress and will be posted here as they are written.