Inferno Creative Studio

Chapter summaries · The Complete Guide to Photorealism

19: The Future

Reading guide to the Epilogue. Written in 2021, read in 2026 - which is what makes it interesting.

Reading: Eran Dinur, The Complete Guide to Photorealism for Visual Effects, Visualization and Games (Focal Press, 2022), Epilogue, pages 210-214. This page is a guide to that chapter, not a substitute for it.

He opens by marking his own homework

In 2017 Dinur wrote an article naming two emerging projects he thought would spearhead visual effects: Lytro's light-field camera, and Foundry's cloud hub Elara, later Athera. Both evaporated shortly afterwards. Lytro ceased operations in 2018; Foundry closed Athera in 2019.

His assessment of his own error is the useful part. He was wrong at the micro level and right at the macro level. Extensive cloud services are now offered by Amazon, Google and Microsoft and are increasingly used by VFX companies for remote work - the trend was right, the specific product was not. Lytro's light-field camera was too expensive and bulky, but less accurate depth-detection technologies are now standard in consumer phone cameras, and he expects precise per-pixel depth detection in professional cameras eventually.

So he flags the whole chapter honestly: this is in no way an accurate or reliable set of predictions - his guess is as good as yours. That framing is worth keeping in mind, and it is worth reading the rest of the chapter as a snapshot of what looked imminent in 2021 rather than as a forecast.

↑ Contents

Real-time rendering

The trend he is most confident about. The potential of real-time rendering is only starting to be exploited, and the rate of advance leaves little doubt about its future role. He writes with Unreal Engine 5 not yet released, noting its promise to eliminate much of what has held real-time back, and reasoning that if global illumination and raytraced reflections can be generated accurately in real time, other computationally heavy features like true subsurface scattering will follow.

The key claim: until recently, choosing interactive real-time rendering inevitably meant lowering the threshold of realism, detail and believability. That is no longer the case.

He is careful about what real-time does not mean. It does not mean CG is created on the fly - modeling, texturing and shading still take time and effort - and it does not eliminate compositing and matte painting in visual effects. What it gives is on-the-fly layout and lighting adjustment, fast prototyping, flexible pre-visualization for film and TV, virtual environments, on-set visualization for directors and cinematographers, reduced need for expensive render farms, and believable photoreal interactive architectural visualization and immersive games and VR.

↑ Contents

Photorealism on the cloud

His example is Microsoft Flight Simulator (2020), which offered something no flight simulator could before: incredibly detailed and accurate environments spanning thousands of miles, populated with thousands of airports and cities. That volume of data cannot possibly be stored on any single consumer system, so Microsoft stores it on the cloud and streams it continuously during gameplay.

The generalisation: with the ability to store and process massive datasets on the cloud, future games and interactive visualizations could offer a quasi-unlimited level of detail and photorealism, because they would no longer depend on limited end-user specs. The technology for transferring pixels and screen actions in real time over the internet already exists - VFX artists use it to operate software remotely on powerful virtual workstations. Price and dependence on fast internet currently limit it to company-level usage, but that is likely to change.

↑ Contents

LED screens and virtual environments

The section most worth reading for this course, because it is the closest thing in the book to what you do.

The Mandalorian was not the first production to pair LED screens with real-time game engine technology, but it elevated the virtual set and set new standards. Rear projection - background footage screened behind the actor and captured in camera - has been used since the 1930s, but motion discrepancies between camera and screened footage, plus lighting mismatches, made scenes feel artificial. Digital compositing and green screens gave more control and allowed tracking the background to the camera, but green screen still suffers lighting mismatches and a slew of edge and integration issues.

In a virtual set workflow, the 3D environments are built in a game engine during pre-production and displayed on giant LED screens that partially or fully surround the set. The advantages:

  • The screen physically lights the set, essentially acting as a real-world sky light - the same role image-based lighting plays in CG rendering. So the lighting on the actors matches the virtual environment surrounding them, automatically.
  • The physical camera is connected to the game engine's virtual camera, so the displayed background continuously reacts to the physical camera's movement, angle and even focal changes. Perspective and parallax are always preserved.
  • Real-time rendering allows on-the-spot changes to layout and lighting within the environment.
  • Because foreground actors, physical set and virtual background are all captured in camera together, there are no green screen extractions, no spill, no edge issues, and potentially no additional VFX work in post.

The limitations he names: LED surround screens are expensive and bulky compared with green screens, and can only be set up in a large indoor soundstage; the virtual environments must be created and completed before shooting starts; and like IBL sky domes, environments on LED screens are merely eggshells - actors cannot walk through them or physically interact with them.

↑ Contents

Machine learning

Written in 2021, and this is where the chapter has aged most. His examples are careful and narrow.

Rotoscoping. Our instinctive ability to identify and separate objects in 2D footage is an extremely difficult problem for computers. The more a machine learning algorithm studies footage of humans, the more accurately it can identify and rotoscope them. He writes that current applications are still far too crude and inaccurate to replace roto artists, but it is logical to assume that separating an actor from a background could eventually be done well.

Upscaling. Online services that enlarge low-resolution photos two, four or eight times, trained by analyzing thousands of images, filling in missing pixels by contextually matching imagery - an automated matte painting of sorts. Will they replace matte painters? Probably not, he writes. But machine learning could plausibly assist artists with time-consuming tasks: searching and collecting reference, matte painting workflow, color-matching in compositing, analyzing footage for lighting, depth-slicing 2D material.

Read this section knowing when it was written. It predates the generative image and video models that arrived over the following years, and his framing - machine learning as an assistant to artists on tedious tasks - is a snapshot of a specific moment rather than a settled position. That does not make it wrong, and the underlying question he raises is still the live one: which parts of this work are craft, and which are labour.

↑ Contents

Procedural environments

Chapter 14 covered procedural modeling for terrains and trees. Here he extends it: algorithm-driven automation generating entire environments populated with terrain, plants, buildings and even live animated creatures and characters.

His honest example is No Man's Sky (2016), released promising unlimited exploration of infinitely many procedurally generated worlds and initially receiving some disappointment, with players criticising boring gameplay and repetitive procedural environments - though the game has since been substantially improved and is now viewed positively.

He notes that while the world's best-selling game, Minecraft, is entirely procedural, few photoreal games are currently based on fully automated purely procedural generation. Most combine pre-built assets with some form of algorithmic distribution. The potential he sees is in systems that generate entire environments and all their assets from scratch based only on a set of physical, thematic and aesthetic rules - useful not just for games but for VFX, visualization and simulators. And the amount of complexity and realism depends entirely on the quality of the algorithm: how closely it emulates the complex relationship between chaos and order in natural environments, or the function-driven expansion of human dwellings.

↑ Contents

Will photorealism disappear?

The best question in the chapter, and the one worth arguing about in class.

The setup: traditional art shifted from figurative to abstract around the turn of the 20th century. Artists lost interest in recreating reality and looked for new avenues of expression. Will photorealism go out of vogue in the same way?

His answer is no, and the reasoning is that the two cases are not parallel. Drawing and painting always served a much broader function than pure art - historical documentation, family portraits, scientific research, journalism, advertisement, entertainment - and for most of those day-to-day uses, photography emerged as a better alternative. Painting did not abandon realism in a vacuum; it was relieved of it.

Photorealism today serves the same broad range of practical purposes, from visualizations to games. It is much less an artistic, aesthetic or stylistic ideal than a functional necessity for immersive experiences. So it is unlikely to lose importance - the opposite, since the immersive experience is bound to be pushed further.

Then the closing thought, which is the sharpest line in the book:

Photorealism as we currently know it is merely an artificial emulation of photography - nothing more than 2D images on screens. Even VR, he points out, is essentially projected moving pictures. Will we ever shift into a new phase of artificial reality that is genuinely three-dimensional and involves the simulation of touch and smell? Such a shift would overtake photorealism the same way photography replaced painting in the late 19th century.

And then the photo comes out of photorealism, because we will no longer be emulating photography. We will be emulating reality.

↑ Contents

Worth arguing about

  • Dinur says photorealism is a functional necessity rather than an aesthetic ideal. Is that true of projection mapping, or is realism a stylistic choice in this field?
  • He argues painting abandoned realism because photography did it better. What, if anything, would relieve digital art of photorealism?
  • Does a projection on a building count as the three-dimensional artificial reality he describes at the end, or is it still just a picture?
  • He was wrong about two specific technologies and right about the trends behind them. Which parts of this chapter look like the trend, and which look like the product?

↑ Contents