The Future of VJing — Four Years On, Did AI Kill Everyone?

The article is written by a human. Mistakes have been edited by AI.
Original predictions made July 2022. Scored and updated August 2026.
In 2022 I asked several VJ groups what they thought the future of VJing was, published my conclusions, and made six specific predictions with dates attached.
Four years is long enough to check. So this is not a new set of guesses — it is the scorecard, plus the honest answer to the question everyone actually asked: can AI replace a live human at a performance?
Short version: no, and the reason is now demonstrable rather than a matter of opinion. But something else got replaced, and I will be direct about that too.
Who this article is for: VJs deciding whether to keep investing in this craft, and anyone who read the 2022 version and wants to know what held up.
What it solves: almost every article about AI and live visuals is written by someone selling AI or someone frightened of it. This one is written against six dated predictions that can be graded.
What you get by the end: the scorecard, what real-time AI on a stage actually looks like in 2026 with hardware figures, what it still cannot do, and where I think this goes to 2030.
Hello dear VJs and clients of LIME ART GROUP. I am new media artist Alexander Kuiava. I made the predictions below, so I get to mark my own homework — which is only useful if I mark it honestly.

The 2022 scorecard

What I predicted in 2022 Verdict
A DALL·E-quality video model in 2–3 years, generating 3D animation and 4K footage Correct, roughly on schedule
VJing is live performance, not content generation — AI cannot do the performance Correct, and now provable
Audio-reactive is the past, not the future Half right
Interactive VJing for 500–10,000 people simultaneously is not possible Still true
AR could change everything, but needs 6G Right about AR, wrong about the bottleneck
Virtual and remote VJing is the gold mine Wrong

Details below, starting with the one everybody argues about.

eye separator

Can AI replace a live VJ?

eye separator
Resolume Arena Software

What AI can do now that it could not in 2022

I predicted a high-quality video equivalent of DALL·E within two to three years. That arrived, and the current state is worth stating precisely, because most people are still working from a 2023 mental model.

Resolution and quality. Kling 3.0 generates natively at 4K, 3840×2160, up to 60fps — the first time a generative video model meets broadcast delivery standards without external upscaling. Google’s Veo 3.1 leads on photorealism and 4K output. Runway Gen-4 delivers the best temporal consistency and motion control currently available.

Clip length. The best models reliably produce 20–25 second clips, with Kling and Veo via Flow reaching up to two minutes of continuous generation. In 2022 we were arguing about four seconds.

Stability. The January 2026 generation of models specifically addressed temporal inconsistency — the old failure where objects shift appearance, colours drift and artefacts appear between frames.

So the prediction was right, and if anything conservative. What it did not tell us is whether any of this touches the live performance.

eye separator
Resolume Arena Software

What real-time AI actually looks like on a stage in 2026

This is the part that matters, and it is genuinely new since 2022.

Generative video is no longer only an offline process. StreamDiffusion reached 91 FPS on a single RTX 4090 using stream batching and stochastic similarity filtering, which proved diffusion models could run live. StreamDiffusionV2, released in November 2025, is a training-free pipeline for interactive live streaming: first frame within 0.5 seconds, 58 FPS with a 14-billion-parameter model and 64 FPS with a 1.3-billion-parameter model — though on four H100 GPUs, which is data-centre hardware, not a laptop in a booth.

At the low-latency end, Decart’s MirageLSD transforms an infinite video stream in real time with under 40 milliseconds of response. Runway announced a real-time system in May 2026 turning a single reference image into a streaming avatar at 24fps HD with about 1.75 seconds end-to-end latency.

And here is the sentence that answers the question in the title. VJs are already performing with this. The documented pattern is real-time diffusion running inside TouchDesigner, with MIDI controllers adjusting prompts and LoRA weights on the fly and beat detection triggering style changes — with the AI functioning as another instrument in the rig.

Read that carefully, because it is the opposite of replacement. A model that responds to a fader is a synthesiser. Somebody still has to play it.

eye separator
Resolume Arena Software

What AI still cannot do

Four limits, all current as of 2026 and none of them close to solved:

It does not read a room. Nothing in a diffusion pipeline knows that the crowd went quiet, that the DJ dropped an unannounced edit, or that the artist walked off the riser eight bars early. The entire craft of VJing is response, and response requires being present.

Temporal stability is unsolved. Maintaining exact textures, lighting and object shape across a long sequence is computationally enormous, and complex scenes with people remain detectable. For a four-second loop this does not matter. For a controlled eight-minute show it does.

No alpha channels, no guaranteed loops, no registration. Generative output arrives as flat frames. Anything that must sit precisely on architecture, or loop seamlessly, or composite over a live camera feed, still needs a person to cut it.

The frontier runs on server hardware. The 91 FPS figure is a single RTX 4090 and that is genuinely usable. The 58 FPS 14B-parameter figure needs four H100s. A festival booth is not a data centre, and the models that fit in a laptop are not the models in the headlines.

eye separator
Resolume Arena Software

The honest answer

AI has not replaced the VJ. It has replaced generic content.

That is the real change since 2022, and it is uncomfortable for a company that sells visual content, so let me say it plainly. The bottom tier of visual production — the anonymous filler clip, the interchangeable abstract loop — no longer has a defensible price. Anyone can generate an acceptable version of it in an afternoon.

What went up in value is everything that requires judgement: live response, coherent visual identity, content registered to a specific surface, and knowing which of forty generated options is the right one. Which is exactly what I argued in 2022 when I said people were confusing VJing with content production. That distinction turned out to be the whole thing.

So the 2022 answer stands: a neural network is not going to stand in front of a thousand people and decide what happens next. But it will absolutely take the work of anyone whose contribution was pressing play on files that could have been made by anybody.

eye separator

The other five predictions, four years later

eye separator

Audio-reactive: half right

I said audio-reactive was the past rather than the future, and pointed at the Winamp visualiser from 2001 to make the point.

The half I got right: it did not become the defining trend. It is table stakes — a competent VJ has reactive presets in the rig and nobody books anyone for having them.

The half I got wrong: I underrated it as a control surface. In the real-time AI setups described above, audio analysis is doing the triggering — beat detection changing style, amplitude driving parameters. Audio-reactivity did not become the show. It became the wiring between the music and everything else, which is more useful than being the show.

eye separator

Interactive at scale: still true

I said you cannot make visuals interactive for 500 to 10,000 people simultaneously, and that a Kinect-based installation only works for small numbers.

This is still correct four years later. The technology for one-to-many personalised interaction at festival scale does not exist, and nothing in the last four years has moved it. The successful applications remain what they were: one performer generating content live, or small-audience gallery and retail installations where individual interaction is genuinely possible.

If you work in galleries, museums or retail, interactive is a real specialism worth building a career on. If you work at festivals, it is still a demo, not a format.

eye separator
Resolume Arena Software

AR: right about the potential, wrong about the bottleneck

I said AR glasses could change the entire industry, and that we would need at least a 6G network to move the data.

The scale arrived faster than I expected. Smart glasses shipments surged 167% year on year in Q1 2026, about 2.25 million units in a single quarter, with IDC forecasting roughly 13.6 million units for full-year 2026 and 27.3 million by 2030. Meta held 69.2% of the market in Q1 on the back of its Ray-Ban partnership, and the Meta Ray-Ban Display launched at $499 in March 2026. Google and Samsung announced Android XR eyewear at Google I/O 2026.

But the capability did not. Look at the split in the numbers: 13.6 million smart glasses, and only about 950,000 AR glasses with an actual display forecast for 2026. And what those displays do is project text, navigation arrows, subtitles and interface elements through waveguide optics — not spatially tracked animated content overlaid on a festival stage.

So my error was the bottleneck. I said the constraint was network bandwidth. The real constraints turned out to be optics, field of view and battery. A glasses-based personal visual layer at a festival is still not close, and it will not be a 6G announcement that changes it.

eye separator

Mobile apps controlling the show: still a mess

I said letting the audience control VJ software from their phones would produce chaos — one person pulling red, another blue, faders fighting each other.

Nothing has happened in four years to suggest otherwise. Where audience phone interaction works, it works as a coordinated single output — the crowd as one pixel grid, a synchronised light effect, a vote — not as individuals steering the visual mix. That is a lighting and crowd-effect technique, not VJing.

eye separator

The one I got wrong: virtual shows

In 2022 I wrote that remote and virtual VJing was “just a big gold mine”, that demand would only grow, that you could earn more for less time, and that any artist anywhere could become your client.

That did not happen, and it is the clearest miss on the list.

Livestreamed shows did not disappear — they are a permanent part of the industry now. But they became a normal, modestly paid format rather than a growth market. Between 2023 and 2025 overall ticket sales in North America flattened or dipped in several categories, festivals that used to sell out in minutes struggled with weekend passes, and mid-level arena tours were scaled back. The market restructured; it did not shift online.

Where I went wrong is instructive. I was reasoning from a pandemic-era spike and treating a constraint-driven behaviour as a preference. When the constraint disappeared, so did most of the behaviour. The lesson I would apply now, including to everything in the forecast below: be careful about extrapolating from a market that is temporarily unable to do the thing it actually wants to do.

What I think happens to 2026–2030

Stated so it can be graded again in four years.

Real-time generative tools become a normal source in the rig, not a replacement for it. By 2030 I expect a real-time model to be as ordinary in a VJ setup as a generator or an audio-reactive preset is now — one more channel in the mix, driven from the same controller. The people who learn to play it will be more employable; the ones waiting for it to go away will not.

The content market splits permanently. Generic clips keep falling toward zero. Coherent sets with a point of view, content registered to specific surfaces, and anything requiring alpha, precise loops or architectural fit hold value. This is already happening and it accelerates.

Live performance value rises, not falls. This is the counterintuitive one, and it follows directly from the above. As the image becomes cheap, the scarce thing becomes the judgement — what to play, when, and why. That is priced as a person, not as a file.

The laptop catches up with the data centre, slowly. The gap between a 4090 running a small model and four H100s running a large one will narrow but not close by 2030. Expect useful real-time generation on portable hardware, at lower parameter counts, with the frontier permanently elsewhere.

AR remains a specialism, not a format. Display-capable AR glasses grow, but a spatially tracked personal visual layer for a festival crowd is a 2030s question at the earliest. The near-term AR work for visual artists is in retail, museums and brand activations, where the audience is small and the hardware can be supplied.

Interactive at true festival scale still does not arrive. I was right in 2022 and I expect to be right again. It is a one-to-many problem that nobody has a mechanism for.

eye separator

What to actually do about it

Learn to drive a real-time model as an instrument, not to prompt one as a content tool. TouchDesigner with a diffusion pipeline, mapped to your existing controller, beat-driven. The skill is the same skill you already have — knowing when to change — applied to a new source.

Stop competing on generic content. Whether you produce it or license it, the differentiator is coherence: one palette, one motion language, a recurring form. A set with a point of view is not replaceable by a prompt; a folder of nice abstract clips is.

Keep the parts of the job a model cannot hold. Reading the room. Being in the room. The relationship with the show director and the lighting designer. Those are not nostalgic virtues, they are the reason a human is still in the booth.

Do not extrapolate from a spike. That is the mistake I made with virtual shows in 2022 and it is the most common error in every “future of” article, including the ones written this year.

Sources

Predictions and opinions are our own. Technical and market figures were checked in August 2026.

Real-time and generative AI

  • StreamDiffusion and StreamDiffusionV2 — published performance figures for real-time diffusion pipelines
  • Decart — MirageLSD live-stream diffusion latency
  • Runway — real-time video system announcement, May 2026
  • Comparative analyses of Kling 3.0, Google Veo 3.1, Runway Gen-4 and Sora, 2026
  • Documented practice of real-time diffusion in TouchDesigner for live visual performance

AR and smart glasses

  • IDC — smart glasses shipment forecasts for 2026 and 2030
  • TrendForce — AR glasses shipment forecast for 2026
  • Q1 2026 smart glasses market share reporting

Live events

  • Reporting on North American ticket sales and festival performance, 2023–2025
  • Analyses of livestream music demand after the pandemic

Our own material

  • LIME ART GROUP community discussion on the future of VJing, 2022, and the resulting predictions
  • Production and performance experience since 2012
eye separator

Grade me again in 2030

The predictions in the forecast section are dated and specific on purpose. If they are wrong I would rather find out from my own follow-up than from a comment.

If you disagree with any of them — particularly the claim that live performance value rises as images get cheaper — the argument is worth having in our VJ communities, where the 2022 version of this article started.

And if you are building a set that a prompt cannot replace, the components are in VJ Loops Packs and AI Visuals — the second of which exists precisely because generated content still needs a person to correct, cut and encode it before it survives a real show.

Be a VJ. Use whatever you have and mix it. Do not be afraid of the models, and remember what it means to be the one in the room.

Latest Feature VJ Loops

Hand-picked visuals by our New Media Artists

Best wishes from Vienna,
Thanks for your attention, faithfully yours,
Alexander Kuiava – Founder & CEO LIME ART GROUP
https://alexanderkuiava.com/