Der Fahrplan · Schritt 5
Images and video with AI
Which AI image, video and audio tools fit which job, what they cost, and how to handle rights and labelling without a lawyer.
Tools in diesem Schritt
Use media where it does a job
The fastest way to make a site look machine-made is to put a generated hero image on every page. Readers now recognise the glossy, over-lit AI look instantly, and it signals that nobody cared. So the first decision is not which tool but whether. An image earns its place when it shows something words cannot: a screenshot of the tool in use, a diagram of a workflow, a chart of real numbers, a photo of the actual thing. Decoration does not earn its place.
With that rule in place, generated media is useful in three situations: a diagram or illustration you cannot photograph, a consistent visual style for a series, and video or audio versions of content you already wrote. Here is what to use for each and what it currently costs. Prices move; check every current page before you subscribe.
Images
Midjourney
Midjourney still produces the most consistently attractive images and has the strongest style controls. It has no free tier; the Basic plan is $10 a month, Standard is $30 with unlimited slower generations. Its weakness is text: it cannot reliably render words in an image. Its risk is sameness; its default look is the AI look everyone recognises, so you have to push it with style references and your own colour rules.
Fits: illustrations for a series, brand imagery, anything that needs to look designed.
Ideogram
Ideogram is the tool for images that contain words: posters, quote cards, thumbnails, social graphics with a headline. Its text rendering is the best of the group. Basic is $8 a month for 400 priority credits, Plus is $20 monthly. Check the current page.
Fits: thumbnails, social cards, anything typographic.
Flux
Flux from Black Forest Labs is a family of models with open weights, which means you can run it through any of a dozen API providers, inside automation tools, or on your own hardware. Quality is close to Midjourney, prompt adherence is often better, and per-image API pricing is a few cents. It is the model to build a pipeline on: if step nine's automations need an image for every article, this is how you generate it without a human clicking.
Fits: automated pipelines, developers, anyone who wants a model rather than a subscription.
Video
Generated video has crossed from novelty to usable for short clips, and it is still expensive per second. Budget for it as a treat, not a habit.
Runway
Runway is the professional's choice: precise camera control, image-to-video, and editing tools around the generator. Standard is $12 a month billed annually for 625 credits, which buys under a minute of its best model; Pro is $28 on the same terms. Costs per second vary by model, so read the credit table.
Fits: a few cinematic seconds for an intro, a product shot in motion, a hero loop.
Kling
Kling produces long, coherent clips with natural motion and is the cheapest way to get usable generated video. Standard is $10 a month for 660 credits, Pro is $37. Check the current page.
Fits: social clips, b-roll for a talking-head video, experiments.
HeyGen and Synthesia
Both make a presenter read your script on camera, in any language, without you filming. The Creator plan at HeyGen is $29 a month ($24 annual) and includes stock avatars and lip-synced translation. The Starter plan at Synthesia is $29 a month ($18 annual) with about ten minutes of video and includes a custom avatar, which HeyGen reserves for its Business tier. Check both pages.
Fits: training videos, explainers, a translated version of a video you already made. Does not fit: anything that trades on you being a real person. An avatar that is clearly an avatar is fine. An avatar pretending to be you at your desk is a trust problem and, in the EU, a labelling obligation.
Audio
ElevenLabs
ElevenLabs is the standard for narration and voice cloning. Starter is from $5 a month with a commercial licence and about thirty minutes of speech; Creator is $22 with professional voice cloning and around a hundred minutes. Starter is the minimum tier if you publish or monetise. A cloned voice of yourself, reading your own articles, is the most honest use: it is your words and your voice, and you can say so.
Fits: article narration, podcast versions, dubbing.
Suno
Suno generates full songs from a description. Commercial use requires the Pro plan, $10 a month; the free tier is for personal use only. Rights over generated music are unsettled in several countries, so treat it as background for your own videos, not as a product to sell.
Fits: intro music, background beds, jingles.
Rights, in plain terms
Three things are true at once and none of them is legal advice:
- Every tool has its own terms about who owns the output and whether you may use it commercially. Free tiers often forbid commercial use. Read the terms of the tool you actually use, on the day you use it.
- Copyright in purely machine-generated images is weak or absent in several jurisdictions. That means you may not be able to stop others from reusing your generated hero image. It also means you should not build a brand asset, like a logo, on a raw generation.
- Generating a likeness of a real person, a recognisable brand, or the style of a living artist is where the real risk sits. Do not.
Labelling: the one paragraph you need
From 2 August 2026, Article 50 of the EU AI Act applies. In practice it means two things for a creator: if you publish AI-generated or manipulated image, audio or video that could pass as real (a synthetic presenter, a cloned voice, a realistic scene that never happened), you must disclose that it is artificially generated; and if you publish AI-generated text on a matter of public interest without real human editorial responsibility, that text needs a label too. Content that is clearly artistic or fictional gets lighter treatment, and material published before that date is not retroactively affected. This is a summary of a law that has guidelines and exceptions, not advice; the European Commission publishes its own AI Act pages and those are the source to read. Our practice is simpler than the law: we label anything a reader might mistake for a photograph or a recording, in the caption, in plain words.
A small workflow that works
- Decide whether the page needs media at all.
- Write the caption first. If the caption says nothing, cut the image.
- Generate at two sizes, compress, and name the file with real words for search.
- Add alt text that describes what is shown, not "AI generated image".
- Label what needs labelling, in the caption.
Checklist
- Removed decorative generated images that do no job
- Chosen one image tool and one video tool and read their commercial terms
- Checked current prices on every vendor page before subscribing
- Added plain-language labels to any media a reader could mistake for real
- Compressed, named and alt-texted every file before upload