Rendering Improvements

At the end of our second year at The Game Assembly we had a course where we were allowed to create almost anything we wanted. I was interested in exploring other rendering techniques than the deferred rendering pipeline we had created for our own engines. During development of our engine I spent a lot of time on the rendering system, during this time I stumbled upon several things that made the rendering code more complex and less streamlined than what I would have liked.

Because of this, I took this project as an opportunity to try out some more modern techniques that not only offer better performance in common demanding situations, but also streamline the rendering pipeline. And so, the goal of this project was clear, I would implement some techniques that would streamline the rendering pipeline and bring down the complexity, remove unnecessary state changes, bring down the required amounts of shader combinations and implement clustered shading.

Clustered Shading

Clustered shading visualization of the scene's light grid.

We had previously implemented deferred shading for opaque geometry to get better performance with our requirement for having tons of lights in the scene. In my research of finding a more modern solution that would offer even better performance I found clustered shading. Since this technique can be heavily parallelized, I wanted to implement the technique through compute shaders. Below is an explanation of how clustered shading works and what steps the engine performs to render a scene with it.

Clustered shading is an extension of tiled shading, one could think of clustered shading as being the 3D version of tiled shading. In tiled shading we split the screen into tiles of a set width and height, this builds a grid over the entire screen, for clustered shading we take this a step further by building a 3D grid of froxels, taking into account the perspective of the camera when building the grid. With the information about the boundaries of these cells in hand we now go through them all and look at the lights in the scene, our goal is to figure out if the volume of the lights intersects with the boundaries of the cell, and if so, we add it to the list of lights in the cell. We should now have information about which lights affect any given cell, we can now render the scene geometry, clustered shading can be combined with either deferred or forward rendering for different pros and cons. When it comes time to shade a fragment we need to figure out which cell it corresponds to, for this we need camera space position of the fragment, since the froxels were built relative to the camera. After we figure out the cell the pixel is inside we can get the list of lights affecting this pixel, and with this information we can now loop through them all and compute the final shaded color.

For my implementation of clustered shading I decided to do all steps through compute shaders. The first step is to compute the froxels in camera space, this is also done through a compute shader, I decided to divide the screen into 16x9x24 chunks, since this matches the aspect ratio our games were designed for, although running it with a different aspect ratio is no problem either. To maximize usage in thread groups when dispatching I balanced the amount of threads to use and the amount of groups to dispatch, taking into consideration that the amount of threads in a group is usually 32 or 64 on modern GPUs, this was done to maximize the amounts of threads used in the groups.

During my research on how to achieve good depth subdivision of the clusters I stumbled upon a presentation from Siggraph 2016 by Tiago Sousa on how DOOM solves this problem, from there I got an algorithm that showed good results.

We need to add all lights intersecting any cell into the list of lights for that cell. This is again done through a compute shader by checking if the AABB of each cluster intersects the light volume, depending on the type of light I have different intersections tests.

Finally, after all these steps we now have all the information needed to render the scene, as stated before we can use either deferred or forward rendering, I went with forward rendering this time, among other things it allows me to apply hardware MSAA. When the time comes to compute the shading of a fragment we compute or retrieve the camera space position, from this we can compute which cluster it corresponds to in the same way as before, from this we can extract the list of lights, looping through all lights and computing their contribution to the current fragment we get the final color.

Skinning with Compute shaders

An annoying problem I faced while developing the rendering system for our engine, which also added extra complexity, is the fact that all animated meshes require extra care when rendering, all vertices need to be transformed by the corresponding weighted bone transforms.

This is not a big problem per se, but it becomes burdensome when all these meshes require a separate vertex shader that applies these bone transforms, something that static meshes do not require. This requires binding an extra buffer for the bone transforms themselves, as well as binding another vertex shader. If we want to allow meshes to have a custom vertex shader, which might be useful for some vertex effects or such, we must also write yet another vertex shader with the same effect that also applies the bone transforms. We store our shaders in configurations, a configuration consisting of a vertex, geometry and fragment shader. Through a material we know which shader configuration to use while rendering a specific mesh, this works nicely until we need to consider that non-static meshes require a separate vertex shader, this can be implemented but adds a lot to the complexity of the system.

Another problem is the fact that every time we need to render these skinned meshes for different passes or effects, we need a separate vertex shader than for static meshes, since the skinned vertex shader version does more work, they cost more to render, and this must be done every time we need to render the mesh during a frame.

The solution that I went with performs the skinning of the vertices in a compute shader prior to rendering anything and writes the result into a separate vertex buffer that will later be used to render these meshes. This eliminates a lot of this complexity, skinned meshes can use the exact same vertex shader as static meshes, we no longer require extra shader combinations, no extra buffers must be bound while rendering skinned meshes, skinned meshes can use the exact same rendering code as static meshes, this also means that they are cheaper to render multiple times per frame because we no longer need extra work in the vertex shader.

To limit the amount of data that our skinning compute shader needs to write and read per skinned mesh I split the vertex data into multiple streams, a stream with all data that needs to be skinned, another stream with the bone weights and ids and lastly everything that does not need to be adjusted. The skinned data only consist of position, normal and tangent.

With this I brought the complexity of the rendering pipeline down even further and could remove a lot of duplicate code and extra cases, tons of vertex shaders and shader combinations could be removed since they are no longer needed, and it is now easier to create custom shader configurations without having to think about skinned meshes.

Shadow-map Atlas

Shadow maps packed into a shared texture atlas.

During the final shading pass we need to have the shadow-map texture for each light that casts shadows available. When a shadow casting light was created we had to allocate a new texture for the shadow-map, this texture had to be allocated with a particular size without much context to how important it was, how far away the player is, neither could it easily react to dynamic changes as the player moving closer or further away. When rendering the shadows we had to re-bind a new shadow-map texture for every light prior to rendering depth, as well as during the final shading pass we need access to all shadow-map textures since we are now rendering all lights at the same time for a given mesh with the clustered shading approach. To get this working we would need to figure out for a given mesh instance which shadow casting lights affect it and bind all those shadow-map textures to the pipeline prior to issuing the draw calls, as well as a method for the shader to figure out which texture corresponds to which light.

This was not feasible and would lower the benefits achieved from clustered shading as well as increase complexity once again. To fix this I decided to implement shadows through an atlas approach. Instead of shadow casting lights needing a shadow-map texture without any context I would now allocate a large texture when the engine starts up, for now I went with 8192x8192. I created a ShadowMapManager that would be responsible for the texture as well as which lights own a particular area of the map. I set up different tiers of resolution Tier0 - Tier6, Tier0 being the largest at 4096x4096 and subsequent tiers halving the resolution. Now when a light is marked as a shadow caster we can reserve an area from the manager at a select tier, we can calculate the tier from context, how important is the light, how intense is it, how close is it to the player, any kind of metric we can think of. If we already have an area and want to change, we can just release it and reserve a new area at a different tier, if the manager can find an open slot in the texture with the correct size, it is reserved, otherwise we can easily skip rendering shadows for that light until a slot becomes available.

With this approach we can dynamically and easily change the resolution of the shadow-map of a light without any new allocations or anything. When rendering depth to the shadow-maps we can just bind the shadow-map texture upfront, we no longer need to re-bind any textures. When doing the final shading pass, once again we bind the whole shadow-map texture, the shader now has the entire shadow-map available through a single texture with no re-bind needed. Now we only need to figure out which area a particular light uses in the shadow-map, this was solved by adding min and max UV coordinates to per-light data, with this information we can now easily sample from the correct part of the texture, this method was easily implemented for both directional- and spotlights.

With this the rendering flow once again became less complex and more streamlined.

Conclusion

I am pleased with the outcome of this project, I brought down the complexity of the rendering pipeline in our engine, eliminated tons of duplicate code and complex cases, brought down the number of shaders and got to experiment with compute shaders as well as implement clustered shading.

I did not have time to do as much as I would have wanted, if I continued this project, I would like to find a neat way of implementing point lights into the shadow-map atlas, also a big performance improvement with shadows that I did not have time for is to only re-render slots of the shadow-map where something has changed, otherwise leave it the same as the previous frame.