Here is the promised write up, it explains how I made the recent videos, how to structure your work, and good ideas I picked up along the way. https://x.com/i/article/2108301504562262016
Anime Harry Potter
This is a writeup for those that wanted to know how I made the videos that got popular during the Malfoid craze.
The core concept to make the videos is simple. Use a Video Generation model, locally or paid and:
> - Attach a reference picture of Malfoid, Potter, and a castle setting
- Write "She bonks him on the head with a comical squish, stars poping around his head."
- ???
- Profit
The real complexity comes after this.
Ref2Video models
In my videos, I've used
- Minimax H3 local with ComfyUI https://docs.comfy.org/tutorials/video/minimax/minimax-h3
- Minimax H3 Max with FAL https://fal.ai/models/minimax/h3-max/reference-to-video
Minimax H3 Max is a finetuned version by FAL that is relatively cheap compared to the heavy weights like Seedance, and it does the job with the right prompting. Look around for new models, new ones pop up literally every week.
Here is what one of the prompts looks like:
Anime sequence, no music
First shot:
Medium shot of Snape, a tall thin man with sallow pale skin, shoulder-length straight black hair, a hooked nose and cold black eyes, wearing a high-collared black coat buttoned to the neck under a black teacher's robe. He sits at a cluttered wooden desk in a dim stone dungeon office lit by green-tinged candlelight, shelves of glass jars behind him. He is writing on parchment with a quill, not looking up, his expression flat and deadpan. He says "惚れ薬、ね。"
Second shot:
Over-the-shoulder reverse shot from behind Snape, showing Draconia standing in front of the desk with her arms crossed, haughty and slightly embarrassed, glancing to the side. She is the blonde girl from the reference illustration, in a school uniform with a green tie and a large black robe. She says "「惚れ」なんて言ってないわ。でも、まあ、そんなところよ。"
Third shot:
Close-up of Snape sighing through his nose, his eyes still on the parchment as the quill keeps moving. He says in a low monotone "どのような症状がある。"
A few notes I made along the way:
Describe with intent, leaving too much uncertainty in a prompt usually means the model will improvise and that's when your quality starts drifting.
Define the setting, the characters positions, and the interactions before when continuing a prompt. It helps avoiding too much continuity errors.
Get help from a text model to write micro-details. You can write your outline in bullet points, and ask a model to expand them to fully fledged prompts that you need to review and adjust. Claude is good at this, but any flagship model will do.
You can check the initial instructions I've used here: https://pastebin.com/ksnxhWPg
How to build the Refs
Real art and photos
You can kitbash artworks, photos, rebuild something from parts and use that as a ref. You can even film yourself just like the *Devil May Cry* team did for their animation reference. A video like those can be interpreted by a Video model in Ref2Image and it will only keep the movements and gestures if you like, it's pretty amazing.
https://www.youtube.com/watch?v=bK8m6PZC-7U
Pure AI generation
There are a lot of Image Generation models, so I wont make an exhaustive list. The important thing is to find one that you understand how to prompt, and generates stable artstyles you like. Relevant image models in 2026 would be:
- Anime: Anima, gpt-image2.5, Midjourney Niji
- Realistic: Krea 2, gpt-image2.5, Flux
If you already know those, I would suggest diving into more advanced concepts such as using ComfyUI with:
- Img2Img to resample an image, as it's a good way to influence a model into a color palette you want and a simple method to redraw a scene that looks already 90% the way you want.
- ControlNets, to really generate what you want by providing a sketch and letting the model extrapolate, making a rough 3D blockout and providing a depth map for the model, etc. There are a lot of steering tools in local diffusion models.
How to avoid the "AI slop" feel
A good concept helps
Proper writing, pacing, resolution, interesting ideas. What constitutes a novel idea can often be just mixing two concepts that shouldn't go together. What makes writing "cringe" is usually:
1. Naive writing and easy resolutions.
One character gives in too easily, or doesn't push back, things happen without any trouble. I think it was one of the South Park writers that once said "If you don't have a BUT between two story beats, you're in trouble". This can only be improved by reading other works, as you start to see the patterns.
2. Not committing to the concept enough.
Happens when you try a serious or goofy idea and dont go all the way. If you write a dramatic summoning, go all the way: earth shatters, gods are watching, scale doesn't make sense anymore. If you write something corny, stack the clichés: their love is overwhelming and embarrassing (for you, but not for them), the side characters reaction is over-the-top, etc.
Never settle for a bad generation.
This is easier said than done, as I imagine most people don't want to sink money into endless re-rolls on video gen models. There are a few solutions.
1. It helps to know how to salvage footage.
Learn After Effects, Davinci Resolve, or a simpler software that still provide simple tools such as masks, freezing on a frame, muting a clip, keeping only the audio of a clip, and a few audio effects (reverb/delay).
That way, you can usually turn a few bad generations into a good one by keeping the visuals of a clip, but only the audio of another one, and thus building a frankenstein clip that does what you wanted, without spending on a new generation.
2. Get to know your models.
Read the prompting documentation, make sure you understand the weaknesses and requirements. For instance Minimax H3 vanilla requires a very strict syntax, and ignoring that does create meh results. You also need to know cinema keywords to steer the results into more specific settings and framing. Knowing all the types of camera shots, scene techniques, etc. does wonders to vary your clips.
https://www.studiobinder.com/blog/ultimate-guide-to-camera-shots/
When generating images, improve your vocabulary and try weird combinations.
All the models are made of Text Model part that encodes the prompt, and this means it understands very niche concepts, but you need to invoke them. Simply asking for 1girl, long straight blonde hair, black and green dress isn't gonna get you something original.
Take the time to read how things are called, read up on terminology of the thousands of woman hairstyles, accessories, fabric, etc., because the models know these. One example I come back to: there are (at least) 100 different types of fabrics https://www.waynearthurgallery.com/types-of-fabrics_fbic/
