A stick figure YouTube channel is generating $13,000 per month after just two months of uploading. Videos have reached 7.6 million, 3 million, and 1.8 million views. Although the channel was created years ago, these stick figure videos have only been uploaded for the past two months and have accumulated 15 million total views. The goal was to recreate this channel using only free artificial intelligence tools—and it was accomplished in under 30 minutes per video. Here is the complete process, including the critical step most people get wrong and a tip for avoiding YouTube monetization problems with AI-generated content.
The first step is identifying what topics to cover. Screenshots of the viral channel's video titles are captured and brought into Claude along with a specific prompt: "These are screenshots from a viral stick figure channel. Analyze the pattern of titles and topics that get the most views and give me 8 to 10 new topic ideas with the same type of curiosity hook—science, psychology, history, human body. Don't repeat exact topics, just give me different angles."
Claude analyzes the pattern and returns topic ideas. The topics tend to be curiosity-driven: science facts, psychological phenomena, historical oddities, and human body mysteries. From the suggestions, one topic is selected: "Why do you yawn when you see someone yawn."
A prompt instructs Claude to write a narrative YouTube script on the chosen topic with specific parameters: conversational and curious tone, strong hook in the first lines, real scientific data explained simply, and a closing that connects with the viewer. No questions as the opening of the script.
Claude produces a complete script. One adjustment is requested: "Give me the script in running text without section titles or tags in brackets, ready to paste directly into a video generator." This produces clean narration text ready for voiceover.
The script goes to ElevenLabs using the free plan, which provides 10,000 credits. Under Voices, the language is set to Spanish, and a conversational voice is selected. The script is pasted, and the voiceover is generated and downloaded.
For those concerned about YouTube's AI content policies: YouTube has primarily removed channels that uploaded mass-produced videos using AI tools that created repetitive, nearly identical content across multiple channels. Traditional faceless channels on YouTube have used AI voices for years without issues. One channel was created years ago with no equipment; some videos were narrated with the creator's own voice, others with AI tools, and there were never any problems. For additional security, the voiceover can be recorded using a phone microphone or by hiring a narrator on Fiverr.
This is the step most people get wrong. These videos have a fundamental structure between the voice and the images. The natural pauses in the narration must be identified before creating any images. This tells exactly how many images are needed and at what precise moment each image should appear.
TurboScript, a free platform that allows three daily transcriptions, is used for this. The voiceover file from ElevenLabs is uploaded, the language is set to Spanish, and the transcription is generated. The resulting transcript includes timestamps marking the natural pauses—at second 07, second 12, second 14, and so on. These timestamps become the exact moments where images change in the final video.
The transcript with timestamps is brought back to Claude along with a reference image of the desired visual style: stick figures, simple sketch-like drawing, hand-drawn appearance. The prompt instructs Claude: "Here is my script with timestamps. I'm attaching a reference image of the visual style I want—stick figures, simple hand-drawn sketch style. Generate an image prompt for each timestamp, replicating that same style—simple strokes, white background, no complex colors, no text in the image. Give me the prompts in order, each associated with its timestamp."
The reference image was created in Gemini. Claude generates the complete list of image prompts, each matched to its corresponding timestamp.
All images are generated for free using Google Flow. A new project is created, and images are selected as the output type. The first prompt from Claude is pasted. Settings are configured to 16:9 format at 1x resolution with Nano Banana selected for image generation. The first image is generated. The process repeats for each prompt until all images corresponding to all timestamps are complete.
Each image is renamed according to its timestamp—for example, "0 to 8" for the image that plays from second zero to second eight. This organization makes the final assembly straightforward.
All images are imported into the video editor and the voiceover is placed on the timeline. The first image covers seconds 0 to 8. The second covers seconds 8 to 14. The third covers seconds 14 to 21. The process continues until every image is placed at its corresponding natural pause. The result is a complete, professionally structured video with narration perfectly synchronized to simple, engaging stick figure visuals.
The final output matches the viral format precisely: curiosity-driven narration, conversational tone, scientific facts explained simply, and stick figure illustrations that change at the natural rhythm of the spoken words.