CHAPTER 3. A SINGLE IMAGE GENERATIVE MODEL FOR TILEABLE TEXTURE SYNTHESIS

Figure 3.1: Results obtained with SeamlessGAN for tileable texture synthesis. We show 2x2 repetitions of the outputs on the squared images, and the input exemplars by their side.

Figure 3.1: Results obtained with SeamlessGAN for tileable texture synthesis. We show 2x2 repetitions of the outputs on the squared images, and the input exemplars by their side.

In this chapter, we present SeamlessGAN, a method capable of automatically generating tileable texture maps from a single input exemplar. In contrast to most existing methods, focused solely on solving the synthesis problem, our work tackles both problems, synthesis and tileability, simultaneously. Our key idea is to realize that tiling a latent space within a generative network trained using adversarial expansion techniques produces outputs with continuity at the seam intersection that can then be turned into tileable images by cropping the central area. Since not every value of the latent space is valid to produce high-quality outputs, we leverage the discriminator as a perceptual error metric capable of identifying artifact-free textures during a sampling process. Further, in contrast to previous work on deep texture synthesis, our model is designed and optimized to work with multi-layered texture representations, enabling textures composed of multiple maps such as albedo, normals, etc. We extensively test our design choices for the network architecture, loss function, and sampling parameters. We show qualitatively and quantitatively that our approach outperforms previous methods and works for textures of different types. The contributions presented in this chapter have led to the following publication1:

“SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps” Carlos Rodriguez-Pardo, Elena Garces

IEEE Transactions on Visualization and Computer Graphics (TVCG)

(2022)

3.1 Introduction

Realistic and high-quality textures are important elements to convey realism in virtual environments. These can be procedurally generated [Tu+20; HDR19; Gue+20; Gal+12; Gil+14; Gui+17; HN18], captured [Guo+20b; LSC18] or synthesized from real images [EF01; Kwa+05; Zho+18; Mor+17]. Frequently, textures are used to efficiently reproduce elements with repetitive patterns (for example, facades, surfaces, or materials) by means of spatially concatenating -or tiling- multiple copies of themselves. Creating tileable textures is a very challenging problem, as it requires a semantic understanding of the repetitive elements, often at multiple scales. For this reason, such a process is frequently done manually by artists in 3D digitization pipelines.

Recent advances in Convolutional Neural Networks (CNNs) and Generative Adversarial Networks (GANs) have been applied to texture synthesis problems [Zho+18; FAW19; Liu+20; BJV17; JBV16; Mar+20; Her+20] showing unprecedented levels of realism and quality, however, the output of these methods is not tileable. Despite recent methods [DH19; BJV17; Rod+19; Mor+17; Li+20; Nik+21] addressing the problem of tileable texture synthesis, we show that they either assume a particular level of regularity or the generated textures lose a significant amount of visual fidelity with respect to the input exemplars. Further, most of these methods have only focused on synthesizing single images. Rendering realistic materials requires more information about their optical properties beyond what represents a single RGB pixel. To this end, it is common to use spatially-varying BRDFs [Des+19], which are optical appearance models parameterized by stacks of images, each one representing a different property, such as albedo, normals, or transparency. As the number of methods that generate texture stacks from physical samples grows [Des+19; DDB20; Guo+20b] so does the need to turn them into tileable texture stacks.

In this chapter, we propose a deep generative model, SeamlessGAN, capable of synthesizing tileable texture stacks from inputs of arbitrary content. In contrast to Wang Tiles [Coh+03], by which a single large texture region is created by concatenating multiple different tiles with matching borders, we aim to automatically obtain a seamless single-tile, which allows for reduced memory consumption and enhanced usability in casual scenarios when the user might lack the necessary artistic skills. Our key idea is to realize that tiling a latent space within a generative network produces outputs with continuity at the seam intersection [FAW19], which can then be turned into tileable images by cropping the central area. Since not every value of the latent space is valid to produce high-quality outputs, we follow a double strategy: First, we train the generative network using an adversarial expansion technique [Zho+18], which provides a latent encoding of the input texture, which can then be decoded into high-quality outputs that double the spatial extent of the input. Second, we use the trained discriminator as a local quality metric in a sampling algorithm. This allows us to find the input of the generative process that produces tileable textures similar to the original, as well as multiple candidates per input exemplar. As opposed to previous work, which focused on maximizing stationarity [Mor+17] and thus might remove important high-level texture features in regular or near-regular textures, our method is focused on maximizing tileability while preserving the original texture as intact as possible, in terms of its stylistic and semantic properties. To allow for the synthesis of stacks of textures, we propose a neural architecture composed of various decoder networks. Despite not explicitly imposing inter-map consistency, we show that it is implicitly guaranteed by how the generative network is trained. Without losing generality, we show texture stacks synthesis of two maps: an albedo and a normal map. We demonstrate that our method outperforms state-of-the-art solutions on tileable texture synthesis of single images and show several examples for synthesizing tileable texture stacks. We validate our design choices through several ablations studies and off-the-shelf perceptual quality metrics.

Figure 3.2: Naïvely tiling the original texture causes discontinuities at the seam intersections as shown in the top row. Our method automatically generates a tileable texture stack from an input exemplar which double the size of the input.

Figure 3.2: Naïvely tiling the original texture causes discontinuities at the seam intersections as shown in the top row. Our method automatically generates a tileable texture stack from an input exemplar which double the size of the input.

3.2 Related work

We review the texture synthesis methods most closely related to our work. For a more comprehensive review, please refer to the surveys in [Akl+18; Raa+18]. We will also mention other related work regarding tileable texture synthesis, and deep internal learning.

3.2.1 Texture Synthesis

Traditionally, non-parametric texture synthesis algorithms worked by ensuring that every patch in the output textures approximates a patch in the input texture. Earlier methods included image quilting [EF01; EL99], GraphCuts [Kwa+03], genetic algorithms [DZP07], and optimization [Kwa+05; PS00]. More recent approaches use variations of PatchMatch [Bar+09; Bar+15] as a way of finding correspondences between generated and input images [Kas+15; Dar+12; Zho+17].

Despite those methods showing high-quality results for textures of different characteristics, recent work on deep parametric texture synthesis shows better generality and scalability, requiring less manual input. Our approach belongs to the category of Parametric texture synthesis. These methods learn statistics from the example textures, which can then be used for the generation of new images that match those statistics. While traditional methods used hand-crafted features [De 97; HB95], recent parametric methods rely on deep neural networks as their parameterization. Activations within latent spaces in pre-trained CNNs have been shown to capture relevant statistics of the style and texture of images [GEB16a; Gat+17; JAF16; Ren+17; RB19]. Textures can be synthesized through this approach by gradient-descent optimization [Sne17; GEB15a] or by training a neural network that learns those features [DB16; Nik+21]. Finding generic patterns that precisely describe the example textures is one of the main challenges in parametric texture synthesis. Features that describe textured images in a generic way are hard to find and they typically require hand-tuning. Generative Adversarial Networks (GANs), which have shown remarkable capabilities in image generation in multiple domains [Kar+18; KLA19; Kar+20b], can learn those features from data. Specifically, in texture synthesis, they have proven successful at generating new samples of textures from a single input image [Zho+18] or from a dataset of images [BJV17; FAW19; JBV16; Liu+20; Mar+20]. We build upon the method of Zhou et al. [Zho+18] which shows good performance on the synthesis of non-stationary single image textures, and extend it to synthesize texture stacks, as well as generate tileable outputs.

3.2.2 Texture Tileability

Synthesizing tileable textures received surprisingly little attention in the literature until recent years. Moritz et al. [Mor+17] propose a non-parametric approach that is able to synthesize textures from a single example while preserving its stationarity, which measures how tileable the texture is. Li et al. [Li+20] propose a GraphCuts-based algorithm. They first find a patch that optimally represents the texture, then use graph cuts to transform its borders to improve its tileability. This method allows for synthesizing multiple maps at the same time. Relatedly, Deliot and Heitz [DH19] propose a histogram-preserving blending operation for patch-based synthesis of tileable textures, particularly suited for stochastic textures. The power of deep neural networks for tileable texture synthesis has also been leveraged in the past years. First, Rodriguez-Pardo et al. [Rod+19] exploit latent spaces in a pre-trained neural network to find the size of the repeating pattern in the input texture. Then, they use perceptual losses for finding the optimal crop of the image such that, when tiled, looks the most similar to the original image. Also leveraging deep perceptual losses and using Neural Cellular Automata [Mor+20] as an image parameterization, Niklasson et al. [Nik+21] generate self-organizing textures that are seamlessly tileable by design, but are limited in resolution and by the quality of the gram matrix perceptual metric used as a loss function [Hei+21]. Bergmann et al. [BJV17] achieve tileability in their output textures by training a multi-image GAN in a periodic spatial manifold. Our proposed method does not follow any of these approaches. Instead, we build upon a state-of-the-art single-image texture synthesis method, which is able to generate high-quality and high-resolution images, and extend it to generate textures which are tileable. To this end, we propose a sampling algorithm that finds the input of the generative process that maximizes a novel tileability metric. Our goal is to preserve the input original texture as intact as possible while imposing artifact-free continuity at the intersection of the seams when the texture is tiled. Furthermore, our method has a reduced computational footprint compared to other deep generative texture synthesis methods, as we show in the results section.

3.2.3 Deep Internal Learning

Learning patterns from a single image has been studied in recent years, in contexts different to those of texture synthesis. Ulyanov et al. [UVL18] show that single images can be represented by randomly initialized CNNs, and show applicability on denoising or image inpainting problems. A similar method is proposed by Shocher et al. [SCI18] for single-image super-resolution. Additionally, single images have shown to be enough for learning low-level features that generalize to multiple problems [ARV19]. Hypernetworks [HDL16] in Implicit Neural Representations [Sit+20; Tan+20] also allow for shift-invariant priors over images, which can then used in generative models [DTD21]. Single-image generative models have been explored for domains other than textures. GANs trained on a single image have been used for image retargeting [Sho+19], deep image analogies [Ben+21], or for learning a single-sample image generative model [SDM19; Hin+21; Sho+19; SGK21; VZH20]. These methods, while powerful for natural images, are not well-behaved for textures, as shown in [Liu+20] and [Mar+20]. By introducing inductive biases specially designed for textured images, characterized by their repeating patterns, deep texture synthesis methods achieve better performance in textured images than generic single-image synthesis approaches, which need to account for more globally coherent semantic patterns.

3.3 Overview

Our method takes as input an untileable texture stack T , which is a layered representation śor SVBRDF [BS12]ś of a material used by render engines to virtually reproduce it. A typical texture stack is composed of several maps such as albedo, normals, or roughness. For simplicity, and without losing generality, we assume that our texture stacks have two maps: an albedo A, and a normal map N, both are RGB images of the same dimensions. There are several methods that generate texture stacks for any material using, for example, smartphones [Des+19; DDB20; Guo+20b; Shi+20; VPS21; Rod+23a], however, these stacks are not tileable by default. Having this data as input, we propose a two-step automatic pipeline to generate tileable texture stacks from arbitrary material input exemplars.

In the first step, described in Section 3.4, we use crops t of the input texture T to train a Generative Adversarial Network (GAN) able to synthesize novel untileable texture stacks T using adversarial expansion [Zho+18]. This training framework has been shown to provide state-of-the-art results on single-sample texture synthesis, surpassing previous approaches [Uly+16; BJV17; Hen+21]. Thanks to using a GAN, we learn an implicit representation of the texture parameterized in two neural modules: G and D.G:tT is a generator that outputs new untileable textures which double the spatial resolution of the input. D is a discriminator that, thanks to the adversarial framework used for training, is able to distinguish real from fake textures.

In the second step, described in Section 3.5, we produce tileable stacks by means of two key ideas: first, we tile a latent space x0 of the trained generator G, obtaining a novel texture stack showing continuity at the seams intersection T^. Second, we implement a sampling process using the trained discriminator D used as quality metric Q to find an optimally tileable texture stack T. An overview of our sampling step is shown in Figure 3.3.

Figure 3.3: Overview of SeamlessGAN. A crop t of the input texture stack T, is fed to an encoder E, which transforms it into a latent space x0E(t). We tile this latent space vertically and horizontally, obtaining a latent field F0. F0 is further processed by several residual blocks Ri,i{1,l}.. The resulting latent variables are transformed by two different decoders, GA and GN, which output 4 copies of a candidate tileable texture T^. By cropping the center part of this texture, we obtain a single of those copies, with seamless borders, T^c. Additionally, this texture can be analyzed by a discriminator D which provides local estimations of the quality of the synthesis. We introduce a tileability evaluation function Q, which, by analyzing two vertical and horizontal search areas sv,sh, is able to detect artifacts that may arise when tiling the texture. This gives us an estimation of how tileable the texture is. This estimation can then be used by a sampling algorithm for generating high-quality tileable textures.

Figure 3.3: Overview of SeamlessGAN. A crop t of the input texture stack T, is fed to an encoder E, which transforms it into a latent space x0←E(t). We tile this latent space vertically and horizontally, obtaining a latent field F0. F0 is further processed by several residual blocks Ri,i∈{1,l}.. The resulting latent variables are transformed by two different decoders, GA and GN, which output 4 copies of a candidate tileable texture T^. By cropping the center part of this texture, we obtain a single of those copies, with seamless borders, T^c. Additionally, this texture can be analyzed by a discriminator D which provides local estimations of the quality of the synthesis. We introduce a tileability evaluation function Q, which, by analyzing two vertical and horizontal search areas sv,sh, is able to detect artifacts that may arise when tiling the texture. This gives us an estimation of how tileable the texture is. This estimation can then be used by a sampling algorithm for generating high-quality tileable textures.

In the following, we first describe our GAN architecture, including our proposal to deal with textures stacks. Then, we describe the sampling process to generate tileable ones.

3.4 Self-Supervised Texture Synthesis with Adversarial Expansion

For each input texture T we train a GAN, whose generator G is able to synthesize novel examples T . Unlike other GAN frameworks, which take as input a random vector, we use crops t of the original stack as input to guide the generative sampling, such that, T^=G(t) for tT. The training strategy builds upon the work of Zhou et al. [Zho+18], which uses adversarial expansion to train the network as follows: First, a target crop tT of 2k × 2k pixels is selected from the input stack. Then, from that target crop t, a source random crop tst is chosen with a resolution of k × k pixels. The goal of the generative network will be to synthesize t given ts. This learning approach is fully self-supervised. The generative model is trained alongside a discriminator D, which learns to predict whether its inputs are the target crops tT or the generated samples T = T˙=G(ts). Figure 3.4 shows an overview of the training strategy.

Figure 3.4: An overview of our training framework for learning to synthesize texture stacks through adversarial expansion. At each iteration, from an input stack T , we randomly crop a target crop t, from which we select a source crop ts. The goal of the generator G is to estimate t given ts. We measure the difference between target and estimated crops using a combination of adversarial, pixel-wise, and perceptual loss functions.

Figure 3.4: An overview of our training framework for learning to synthesize texture stacks through adversarial expansion. At each iteration, from an input stack T , we randomly crop a target crop t, from which we select a source crop ts. The goal of the generator G is to estimate t given ts. We measure the difference between target and estimated crops using a combination of adversarial, pixel-wise, and perceptual loss functions.

3.4.1 Network Architecture

SeamlessGAN is comprised of an encoder-decoder convolutional generator G with residual connections, and a convolutional discriminator D. So as to be able to synthesize textures of multiple different sizes, the networks are designed to be fully-convolutional. We follow the residual architectural design in [Zho+18], with two extensions: First, building on recent advances in style transfer algorithms, we use Instance Normalization [UVL16] before each ReLU non-linearity in the generator and each Leaky-ReLU [MHN13] operation in the discriminator. This allows to use normalization for training the networks without the typical artifacts caused by Batch Normalization [Cho+18; Des+18; NK18; Kur+19]. Second, in order to allow for the synthesis of multiple texture maps at the same time, we propose a variation in the generator architecture. In the following, we describe the details of each component, with a particular focus on the elements that are different from [Zho+18].

Generator

The generative architecture proposed is comprised of three main components: An encoder E, which compresses the information of the input texture t into a latent space, x0E(t), with half the spatial resolution of the input texture. Then, a set of residual blocks, Ri,i{1,l},, which learn a compact representation of the input texture xiRi(xi1)+xi1. Finally, a stack of decoders G that transforms the output of the last residual block Rl into an output texture TG(xl). Residual learning allows for training deeper models with higher levels of visual abstraction and which generate sharper images [He+16; RFB15; Iso+17]. Similar residual generators have been used in unsupervised image-to-image translation problems [Zhu+17a].

Synthesizing a texture stack of multiple maps with a single generative model poses extra challenges over the single map case. Each map represents different properties of the surface, such as geometry or color, resulting in visually and statistically different images. This suggests that independent generative models per map could be needed. However, the texture maps must share pixel-wise coherence, which is not achievable if multiple generative models are used. Inspired by previous work on intrinsic images [Jan+17; YS19; LS18], we propose the use of a generative model that learns a shared representation of all the texture maps, but has different decoders for each of them. The latent space within the generator will thus encode a high-level representation of the texture, which is then decoded in different ways for each different map in the texture stack. Specifically, we train a stack of decoders, G={GA,GN}, one for each map of the stack (albedo and normals in our case).

Discriminator

For the discriminator D, we use a PatchGAN architecture [Zhu+17a; Iso+17; Led+17; Zho+18], which, instead of providing a single estimation of the probability of the whole image being real, it classifies this probability for small patches of it. This architecture has several advantages for our problem. First, it provides local assessments of the quality of the synthesized textures, which we exploit for obtaining high-quality textures. Second, its architecture design allows to provide some control over what kind of features are learned: by adding more layers to D, the generated textures are typically of a higher semantic quality but local details may be lost. A comprehensive study on the impact of the depth of D can be found in [Zho+18].

Loss Function

We train the networks following a standard GAN framework [Goo+14b]. We iterate between training D with a real sample texture stack T and a generated sample T . The adversarial loss Ladv is extended with three extra loss terms: L1a,L1n,, and Lstyle; corresponding respectively to the pixel-wise distance between the generated and target albedo, normals, and a perceptual loss. The perceptual loss Lstyle is computed as the sum of the perceptual losses between target and generated albedo maps, and target and generated normal maps. We follow the Gram Loss described in [GEB16b] as our perceptual loss. We weight the total style loss by weighting different layers in the same way described in [Rod+19; Zho+18; GEB16b]. Our global loss function thus is defined as:

L=λadvLadv+λL1aL1a+λL1nL1n+λstyleLstyle                  (3.1)

3.5 Tileable Texture Stack Sampling

3.5.1 Latent Space Tiling

After training, the generator G is able to synthesize novel samples of the texture given a small exemplar of it. Although these novel samples duplicate the size of the input, they are not tileable by default. Previous work [FAW19] showed that by spatially concatenating different latent spaces in a ProGAN [Kar+18] generator, it is possible to generate textures that contain the visual information of those tiles while seamlessly transitioning between the generated tiles. Inspired by this idea, we spatially repeat (horizontally and vertically) the first latent space x0 within the generative model, obtaining a latent field F0. This field is passed through the residual layers Rl and the decoders G to get a texture stack, T^, that contains four copies of the same texture with seamless transitions between them (see Figure 3.3). At first, the latent field F0 shows strong discontinuities at the borders between the four copies. Later, as the texture is passed through the network, these artifacts are progressively transformed into seamless borders by the rest of the residual blocks and the decoder of the model. The intuition behind this idea is that, after training the network with the target texture, each of these latent spaces encode low resolution versions of the input in the spatial domain and semantically-rich information in the deeper layers.

A seamlessly tileable texture stack T^c is obtained by cropping the central region (with an area of 50% of T^) . The predicted texture T^ has 4 × 4 the resolution of the input crop t. Thus, after the cropping operation, T^c has twice the resolution of the input t.

A key parameter to select is the level l at which to split the generative process. As shown in Figure 3.5, generating the latent field F by tiling earlier levels of the latent space forces the network to transform F more times, thus resulting in more seamlessly tileable textures. We confirm the results found in [FAW19] and noticed that generating this latent field at earlier levels of the latent space (l{0,1}) yields the best visual results, by allowing for smoother and more semantically coherent transitions between tiles. We thus tile the output of the encoder x0=E(t), before transforming it by the residual blocks Ri.

Figure 3.5: Impact of the latent field level l on the quality of the output textures T^ = G(t), using the same input t. Generating the latent field Fθ early layers (l = 0) generates the best visual results, later layers either create small artifacts (l = 3) or generate unrealistic textures (l = 5). A zoom of the central crop is shown as an inset.

Figure 3.5: Impact of the latent field level l on the quality of the output textures T^ = G(t), using the same input t. Generating the latent field Fθ early layers (l = 0) generates the best visual results, later layers either create small artifacts (l = 3) or generate unrealistic textures (l = 5). A zoom of the central crop is shown as an inset.

3.5.2 Discriminator-guided Sampling

Our strategy to tile the latent space guarantees that the generated texture is a continuous function with smooth transitions between the boundaries of the tiles. However, in contrast to the algorithm in [FAW19], where the latent spaces are drawn from random vectors, ours are encoded representations of input textures. Thus, the selection of the input that the network receives plays an important role on the quality of the output textures. As shown in Figure 3.6, not all the generated textures are equally valid. This selection can be posed as an optimization problem: t=argmaxtQ(G(t)), where the goal is to find the crop t* that maximizes the quality Q of the generated texture G(t)=T^c. To solve this optimization problem, one option could be to pose it as an optimization in the generative latent space. However, we found that a simpler solution based on sampling already provides satisfactory results. The remainder of this section explains the sampling process and the quality metric.

Figure 3.6: Outputs of D for textures of different qualities. (a) The generated albedos T^1 are artifact-free on the search areas (b), thus the discriminator D(T^1) is fooled to believe T^2 cannot fool the discriminator, which finds artifacts on the central areas D(T^2) (d).

Figure 3.6: Outputs of D for textures of different qualities. (a) The generated albedos T^1 are artifact-free on the search areas (b), thus the discriminator D(T^1) is fooled to believe T^2 cannot fool the discriminator, which finds artifacts on the central areas D(T^2) (d).

Sampling By using a fully-convolutional GAN, our model can generate textures of any size, hardware being the only limiting factor. This is key for tileable texture synthesis as, even if a given texture is seamlessly tileable, larger textures require fewer repetitions to cover the same spatial area, which ultimately results in fewer repeating artifacts, as illustrated in Figure 3.7 (a). There are two main challenges when finding tileable textures: the input texture needs to contain the distribution of the existing repeating patterns, and the input tile itself must not create strong artifacts when tiling the latent spaces of the generator.

Figure 3.7: SeamlessGAN can generate multiple tileable outputs from the same sample T. By cropping different parts of T , we can feed the generator G with different inputs tc (shown on the left), generating tileable textures T^c of different sizes. On the example at the top (a), the synthesized images are tiled so they cover the same spatial extent, which shows that, even if the textures are tileable, larger textures generate fewer repeating artifacts. (b) Shows the tile resulting from sampling different parts of the input image.

Figure 3.7: SeamlessGAN can generate multiple tileable outputs from the same sample T. By cropping different parts of T , we can feed the generator G with different inputs tc (shown on the left), generating tileable textures T^c of different sizes. On the example at the top (a), the synthesized images are tiled so they cover the same spatial extent, which shows that, even if the textures are tileable, larger textures generate fewer repeating artifacts. (b) Shows the tile resulting from sampling different parts of the input image.

Our goal is thus to find the largest possible tileable texture stack. To do so, we sample multiple candidate crops for a given crop size c{cmin,cmax} is the resolution of the input T. We sample crop sizes starting at the largest possible size c=cmax, and stop when we find a suitable candidate according to the tileability metric Q. As shown in Figure 3.7 (b), this sampling mechanism also allows us to generate multiple tileable candidates for a single exemplar if we choose different parts of T as input.

Discriminator-guided Quality Function, Q The second component of our sampling strategy is the quality function used to determine whether a stack is tileable or not. We observed that the artifacts appear on vertical and horizontal frames around the center of the textures (Figure 3.6). This is likely caused by strong discontinuities or gradients on the same areas of the tiled latent spaces, which the rest of the generative network fails to transform into realistic textures. Following recent work on generative models [SSK20], we use the discriminator D as a semantic-aware error metric that can be exploited for detecting local artifacts in the generated textures. This can be done in our case because the global loss function contains pixel-wise, style, and adversarial losses. The adversarial loss learns the semantics of the texture, whereas the L1 and style losses model color distances or repeated patterns [Rod+19]. This combination of loss functions allows us to balance textural, semantic, and perceptual color properties.

We thus design a quality evaluation function Q that estimates if the generated texture stack T^c is tileable, looking for artifacts on a central area, SD(T^c) of the discriminator. This area is composed of two regions, S=svsh, where Sv is a vertical area, and Sh is an horizontal one, both centered on the output of the discriminator (Figure 3.3). The function Q leverages the fact that D outputs 0 when it believes a patch to be synthetic. However, as the values are sample-dependent, we establish a threshold τ using the values of the rest of the image S as a reference τ=γmin(Sr), where Sr=D(T^c)S¯ is the remaining part of the image, and γ is a threshold that allows to control the sensitiveness of Q. Consequently, Q(T^c) is 1 if min(sv) and min(sh) are greater or equal than τ, considering the texture as tileable, and 0 otherwise. The goal of Q is to estimate whether or not the central areas, where discontinuities may be present, are as realistic as the rest of the texture. It works as a classification function which returns a high value if the texture is identified as real by the discriminator in such areas. This allows us to create a sampling strategy that distinguishes between images with local artifacts from seamless textures. By using minimum values instead of the mean estimation of the discriminator, our quality metric focuses on detecting artifacts which may arise during the latent space tiling operation.

3.6 Implementation Details

As described in Section 3.4, we follow a standard GAN training framework, by iterating between training the discriminator and the generator. We use a batch size of 1, and an input size of k = 128. All weights are initialized by sampling a Gaussian distribution N (0, 0.02), following standard practice [Zhu+17a]. G has l = 5 residual blocks with ReLU activations, and D is comprised of 5 convolutional layers with Leaky-ReLU non-linearities, and a Sigmoid operation at the end. We use a stride of 2 for the downsampling operations in E and transposed convolutions [LSD15] for upsampling in the decoders G. We weight each part of the loss function as: λadv=λstyle=1,andλL1a=λL1n=10..

The networks are trained for 50000 iterations using Adam [KB15], with an initial learning rate of 0.0002, which is divided by 5, after iterations 30000 and 40000. Aside from random cropping, we do not use any other type of data augmentation method. The models are trained and evaluated using a single NVIDIA GeForce GTX 1080Ti. Even if training takes around 40 minutes for each texture stack, once trained, the generator can generate individual new samples in milliseconds, before checking for tileability. We use PyTorch [Pas+19] as our learning framework. To accelerate the training process, we leverage mixed precision training and automatic gradient scaling [Mic+18]. Every operation in the training pipeline is done natively in GPU using Torchvision [MR10]. These optimizations allow us to train the networks one order of magnitude faster than previous methods [Zho+18]. The input textures T tested in this chapter have, on average, 500 pixels in their larger dimension, for which our method can generate tiles of at most 1000 pixels. We use Photometric Stereo [Ike81] for computing the normals of fabric textures, and artist-generated normals for the other samples.

We tile the latent space at its earliest level (Fl,l=0), as it provides the best quality results, as shown in Section 3.5.1. For identifying the tileability of textures, we use a search area S that spans a 20% of each spatial dimension of the textures. For all the results shown, γ = 1. We choose cmin=100 pixels , and cmax equal to the resolution of the whole input texture. This means that for the level of cmax, only one texture is sampled. For each c<cmax, we sample 3 different random crops. Typically, for the results shown in this chapter, a tileable texture stack is found at high-resolution crop sizes, ccmax10, thus needing to sample and evaluate less than thirty random crops before finding a satisfactory solution. This entire sampling procedure takes less than five minutes.

3.7 Experiments

In this section, we evaluate two of our main contributions: first, in Section 3.7.1 we study the impact of the design choices of the generator G to synthesize a multi-layer texture stack, and second, the quality of the tileable texture synthesis through several ablation studies. We study the inter-map consistency in Section 3.7.2, and the impact of the loss function in Section 3.7.3. Finally, we compare our method with methods on tileable texture synthesis (Section 3.8).

3.7.1 Generator G Design

One of the main challenges when synthesizing texture stacks using a single generator is to preserve the low-level details of each of the texture maps whilst maintaining the local spatial coherence between them; if this coherence is lost, renders that use the synthesized maps will show artifacts or look unrealistic. We propose two different variations of the single-map generative architecture G presented in [Zho+18], each of which makes different assumptions on how the synthesis should be learned taking into account the particular semantics and purpose of each map. A diagram of each proposed architecture is shown in Figure 3.10. For a fair evaluation, we follow the criteria that both networks must have approximately the same number of trainable parameters.

Our baseline, G1, treats the texture stack as a multiple-channel input image, and entangles every texture map in the same layers. It assumes that the maps in the stack share most of the structural information and, as such, there is no need to generate them separately. Thus, the last layer in the decoder outputs every texture map. Our proposed alternative architecture, G2 finds a shared representation of each texture map, but has a separate decoder for each of them. As such, the residual blocks are shared for all the texture stack, but each decoder can be optimized for the semantics and statistics of each particular map.

To quantitatively evaluate which architecture produces the highest quality output we compare the original texture with the generated one using standard metrics: SSIM, Si-FID, and LPIPS. The Structural Similarity Index Measure (SSIM) [Wan+04] is a perceptual-aware metric, working on the pixel space, that measures the similarity in structural information, and may be appropriate to evaluate synthesized textures. The Si-FID [SDM19] is a single image extension of the Fréchet Inception Score (FID) [Heu+17], which measures the difference in deep latent statistics between natural and artificially generated images. Finally, we use the Learned Perceptual Image Patch Similarity (LPIPS) [Zha+18] as a perceptual distance metric in deep image spaces. This metric is widely used for evaluating generative models [Kar+20b; Hua+18b; Cha+19; Alm+18; Mar+20].

The baseline G1 shows more artifacts than G2, most likely due to the fact that the generation parts of the network are not fully separated. A quantitative evaluation is shown in Table 3.1, showing that G2 outperforms G1 from perceptual and statistical standpoints. Those metrics require the input and target images to have the same spatial dimensions. To obtain this, we crop the 50% center area of each generated stack, which doubles the dimensions of the inputs due to the adversarial expansion approach we follow, and compare them with the input textures. In summary, G2 allows to generate textures that better preserve the properties of the input images without any additional computational cost and without requiring to modify the loss function, or the discriminator design.

Table 3.1: Quantitative comparison between our variations of the generator G. We show the average results on different metrics across different texture stacks, separated by maps. As shown, G2 outperforms its baseline across metrics and texture maps. G1 only yields better scores at the Si-FID metric on the normal map. Higher is better for SSIM, while lower is better for Si-FID and LPIPS

SSIM [Wan+04] ↑

Si-FID [SDM19] ↓

LPIPS [Zha+18] ↓

G1

Albedo

0.2774

0.3462

0.4826

Normals

0.2364

0.3636

0.4870

G2

Albedo

0.3123

0.3275

0.4377

Normals

0.2921

0.4019

0.4340

3.7.2 Inter-map Consistency

Whilst our generators G can generate high-quality tileable pairs of albedo and normal maps, there is no guarantee that those maps are pixel-wise coherent, as no part in the loss function explicitly accounts for this relationship. The Lstyle loss is computed separately for each map in the texture stack, which may generate non-coherent gradients. Furthermore, our architecture of choice G2 separates the decoding of each map in different parts of its architecture, which may hinder the generation of spatially-coherent maps. Nevertheless, we show that this is not a problem in practice. Figure 3.8 shows crops of our synthesized texture maps. It can be seen that pixel-wise coherency between maps is preserved even for challenging geometric structures. We argue that the role of the discriminator is key to detect inter-layer inconsistencies by yielding lower probabilities for non-coherent maps. Even if separating the decoders may increase the risk of incoherent texture maps, this coherence is forced by the discriminator during training. To test this empirically, for the textures in Figure 3.8, we fed the discriminator with a crop of the original stack, a crop with the normals translated 5 px, and translated 100 px, for which we obtain average values of 0.99, 0.35, and 0.32, respectively. This suggests that, since during training, the discriminator receives the entire texture stack at once, it learns to identify spatial inconsistencies between maps.

Figure 3.8: Crops of synthesized textures using our method. We observe that pixel-wise coherence between maps is preserved.

Figure 3.8: Crops of synthesized textures using our method. We observe that pixel-wise coherence between maps is preserved.

3.7.3 Loss Function Ablation Study

As described in Section 3.4, the generator G is trained to minimize a loss function, which is comprised of adversarial, perceptual and pixel-wise components. Each of these have different impacts on the output synthesis. We hereby study the impact of each of these components.

To better isolate the impact of each metric, we perform this study using a limited generator that only outputs albedo maps. As we show in Figure 3.9, the adversarial loss Ladv provides high-level semantic consistency, the L1 norm acts as a regularization method which removes artifacts, while the perceptual loss Lstyle adds additional details to fully represent the texture. Similar findings are reported in [Zho+18]. None of these loss functions yield compelling or artifact-free results on isolation, being the adversarial loss the most important factor of the global function.

Figure 3.9: Ablation study on the impact of the loss function on the quality of the synthesized textures. We evaluate each network using the same input, marked using a blue box. As shown, training G using the full loss function yields the best results. Each component is weighted by a λ, as specified in Section 3.6.

Figure 3.9: Ablation study on the impact of the loss function on the quality of the synthesized textures. We evaluate each network using the same input, marked using a blue box. As shown, training G using the full loss function yields the best results. Each component is weighted by a λ, as specified in Section 3.6.

3.8 Results and Comparisons

Most tileable texture synthesis methods have not shown results in synthesizing texture stacks. While adding this capability is reasonably easy for some non parametric methods [Li+20; Rod+19], others require major changes in their design. In particular, as we have shown in Section 3.7.1, expanding the architecture of deep learning methods for generating more than one map requires special attention to balance the map’s interdependence with the network’s capacity and expected quality. Therefore, in this section, in order to be able to compare our method with other works, we use a reduced version of our generator that only outputs a single albedo map.

Figure 3.10: A comparison of the results of our proposed architectures. Separating the decoders of the network for each map (G2) outperforms the joint-decoder baseline (G1) for both texture maps in the stacks. It generally allows for more varied albedo maps, with a more correct structure and style; as well as more accurate and sharper normal maps. The outputs in this figure are not tiled, for improved visibility of the artifacts created by G1. In blue, we show an up-sampled crop of the output textures.

Figure 3.10: A comparison of the results of our proposed architectures. Separating the decoders of the network for each map (G2) outperforms the joint-decoder baseline (G1) for both texture maps in the stacks. It generally allows for more varied albedo maps, with a more correct structure and style; as well as more accurate and sharper normal maps. The outputs in this figure are not tiled, for improved visibility of the artifacts created by G1. In blue, we show an up-sampled crop of the output textures.

Qualitative Analysis First, we compare with the Texture Stationarization algorithm proposed in [Mor+17] using examples taken from their own dataset, as their code is not publicly available. A key difference between both methods is that, while their method aims to maximize the stationarity properties of the textures using external metrics, our method learns them from the texture itself, using a self-supervised approach. This results in our method better preserving the content of the original input. As shown in Figure 3.11, our method shows compelling results for the kind of textures shown in their paper. In the fence example, our methods better preserves the vertical straight lines. For the wall example, both methods shows compelling results, capturing a different repetition pattern.

Figure 3.11: A comparison with the work on Texture Stationarization by Moritz et al. [Mor+17], using textures from their dataset. The results are tiled 2 × 2 times to help visualization.

Figure 3.11: A comparison with the work on Texture Stationarization by Moritz et al. [Mor+17], using textures from their dataset. The results are tiled 2 × 2 times to help visualization.

As Moritz et al.’s dataset contains mainly textures of human-made environments, we gather a different set of images of greater variety in their regularity and content. Figure 3.12 shows a comprehensive comparison with the other methods. The work by Li et al. [Li+20], based on Graph-Cuts, works reasonably well if the transformation can be done locally at the seams but fails when the required changes are global, as happens in the basket. The Repeated Pattern Detection algorithm proposed by Rodriguez-Pardo et al. [Rod+19] is not able to handle many of these challenging cases, which do not represent grid-like textures, as in the bananas or stars. The Periodic Spatial GAN (PSGAN) proposed by Bermann et al. [BJV17] generates artifacts, not maintaining the content of the texture, as in the daisies or the hieroglyph examples. A similar effect is observable in the method of [Nik+21], which uses self-organizing representations through neural cellular automata. The output of the method is seamless but the integrity of the original texture is not preserved in many cases, for example, in the hieroglyph or the stars.

Figure 3.12: Comparison with previous methods. On the left column, we show the input textures. From left to right, the synthesized results of Histogram Blending, by Deloit et al. [DH19], the GraphCuts algorithm by Li et al. [Li+20], Repeated Pattern Detection by Rodriguez-Pardo et al. [Rod+19], PSGAN by Bergmann et al. [BJV17], Self-Organizing Textures by Niklasson et al. [Nik+21]; and ours. Outputs are tiled a similar number of times (at least twice in each dimension) for better visualization. Our method generally captures better the overall structure of the texture, while providing seamless and seman-tically coherent borders, for enhancing tileability. From top to bottom, we sample c = 4, 12, 5, 17, 13, 1, 19 and 43 input crops before obtaining a tileable texture.

Figure 3.12: Comparison with previous methods. On the left column, we show the input textures. From left to right, the synthesized results of Histogram Blending, by Deloit et al. [DH19], the GraphCuts algorithm by Li et al. [Li+20], Repeated Pattern Detection by Rodriguez-Pardo et al. [Rod+19], PSGAN by Bergmann et al. [BJV17], Self-Organizing Textures by Niklasson et al. [Nik+21]; and ours. Outputs are tiled a similar number of times (at least twice in each dimension) for better visualization. Our method generally captures better the overall structure of the texture, while providing seamless and seman-tically coherent borders, for enhancing tileability. From top to bottom, we sample c = 4, 12, 5, 17, 13, 1, 19 and 43 input crops before obtaining a tileable texture.

Our method favors keeping high-level and bigger semantic structures of the textures, resulting in larger samples with more variety. The modification applied to the texture is the minimum required to make it seamlessly tileable. It provides high-quality outputs for regular textures with uneven illumination and perspective distortion, such as the basket, textures which are greatly irregular, such as the bananas, and stochastic ones, like the strawberry. While non-parametric synthesis methods are typically computationally cheap, obtaining results in less than one minute [Li+20; Rod+19; DH19], at the cost of quality and generality, parametric methods are more expensive and slower. Our method is competitive in computational cost when compared to other parametric synthesis methods. Our entire training and sampling process takes less than 45 minutes. In comparison, PSGAN [BJV17] needs 4 hours to train, the work by Zhou et al. [Zho+18] requires 6 hours, and the method of Niklasson et al. [Nik+21] takes 35 minutes to generate textures limited to 200 × 200 pixels. On the contrary, our method is not limited by the size of the input texture and can synthesize tiles of any size thanks to the fully-convolutional architecture, hardware being the only limiting factor.

Quantitative Comparison Following a similar evaluation scheme as proposed in Section 3.7.1 for measuring the differences between input and synthesized images, a quantitative evaluation is shown in Table 3.2. We use Moritz’s dataset for a fair comparison with every method. The metrics we use for this experiment require that the input and synthesized images have the same resolution. To achieve this, and, in order to account for artifacts in the borders of the generated textures, we tile the synthesized images until they cover the same resolution as their corresponding inputs.

Table 3.2: Quantitative comparison between different methods. We show the average results on different perceptual metrics across a variety of textures.As shown, SeamlessGAN consistently outperforms its counterparts in every studied metric. We tile the outputs until they match the spatial resolution of the input examples. Higher is better for SSIM, while lower is better for Si-FID and LPIPS. We use a color code to highlight best and worst cases.

SSIM [Wan+04] ↑

Si-FID [SDM19] ↓

LPIPS [Zha+18] ↓

Deloit et al. [DH19]

0.1424

1.3471

0.6207

Li et al. [Li+20]

0.2086

0.9529

0.5818

Moritz et al. [Mor+17]

0.1968

0.7620

0.5171

Rodriguez-Pardo et al. [Rod+19]

0.2144

1.2958

0.5137

Bergmann et al. [BJV17]

0.1723

1.4355

0.5624

Niklasson et al. [Nik+21]

0.1753

1.3171

0.5328

SeamlessGAN

0.2341

0.6311

0.4792

Interestingly, methods based on patches [RB19; Mor+17; Li+20] obtain better SSIM and Si-FID scores than previous deep learning-based methods [BJV17; Nik+21]. This difference is not seen in the LPIPS metric. This suggests that patch-based methods preserve better the structure of the input textures than previous neural parametric models. Our method, by combining a variety of loss functions, allows for better preservation of the style and semantic content of the generated textures. Furthermore, our latent space manipulation algorithm allows for seamless borders between tiles, outperforming both previous non-parametric and parametric methods in perceptual and structural similarity. The magnitude of the Si-FID metric varies significantly between the values obtained in Table 3.1 and Table 3.2, which indicates that this metric may be overly sensitive to the global statistics of the dataset.

Limitations and Discussion Our method is inherently limited by the capabilities of the adversarial expansion technique to learn the implicit structure of the given texture. That is, if the input does not show enough regularity, the adversarial expansion fails, as shown in Figure 3.13. It is interesting to see that the synthesis of the dotted texture fails to reproduce the larger dots, as they are very scarce. However, the remaining structure is very well represented.

Figure 3.13: A limitation of our texture synthesis algorithm: Left: Input texture stacks; right: synthesized stacks. The network fails to replicate the pattern if the occurrence is not frequent enough, as we can see in the base color of these two examples. On the other hand, the synthesized normal map is consistent as its repetitive structure is seen frequently by the network.

Figure 3.13: A limitation of our texture synthesis algorithm: Left: Input texture stacks; right: synthesized stacks. The network fails to replicate the pattern if the occurrence is not frequent enough, as we can see in the base color of these two examples. On the other hand, the synthesized normal map is consistent as its repetitive structure is seen frequently by the network.

Tiling single texture maps, even if they contain seamless borders, may generate perceptible repetitions. This is an intrinsic limitation of any single sample tileable texture synthesis. Approaches such as Wang Tiles [Wan+04] can tackle these limitations, but have additional disadvantages like increased memory and run-time consumption, rendering them less practical for real-time applications. Alternatively, procedural methods [Gue+20] are the most effective way to generate material samples preserving textural properties at multiple scales; however, present several limitations in the range of materials that can be generated to make them fully usable.

3.9 Conclusions and Future Work

In this chapter, we have proposed a deep parametric texture synthesis framework capable of synthesizing textures into tileable single-tiles, by combining recent advances in deep texture synthesis, adversarial neural networks, and latent spaces manipulation. Our results show that our method can generate visually pleasing results for images with different levels of regularity and homogeneity. This work is the first method capable of exploiting properties of deep latent spaces within neural networks for generating seamless textures, and opens the opportunity for end-to-end tileable texture synthesis methods without the need for manual input. Comparisons with previous state-of-the-art methods show that our method provides results which better maintain the semantic properties of the textures, while being able to synthesize multiple maps at the same time.

Our method can be improved in several ways. First, the adversarial expansion framework, while powerful, it has some potential pitfalls that hinder its widespread applicability. The same neural architecture is used for every texture but, as discussed, different choices on the architecture make different assumptions on the nature of the textures. We have proposed a generic architecture that works well for many examples, but recent advances in Neural Architecture Search [Mel+21] may provide better priors on the optimal neural architecture to use for each sample. Furthermore, each texture synthesis network is trained from scratch. This is not only computationally costly, but learning to synthesize one texture may help in the synthesis of other textures, as shown by [Liu+20]. Fine-tuning pre-trained models for generating new textures may provide cheaper syntheses. Besides, our discriminator, while capable of detecting local artifacts, provides little control for separating such artifacts and global semantic errors. Recent work on image synthesis may provide guidance onto designing better discriminative models [SSK20], training procedures [Yun+19; Sin+21], image parametrizations [Wan+20a; HZ21; Nik+21] or perceptual loss functions [Hei+21].

Second, we proposed a synthesis solution based on manipulating latent spaces within the generative model, but explicitly training the network to generate tileable textures may provide better results than our approach. Besides, our sampling procedure could be extended for better selection of textures, by comparing histograms of the activations of the discriminator on the selected central area, instead of simply comparing minimum values.

Finally, our method has the advantage of being fully automatic, however, pre-processing the texture images so they are more easily tileable can help the synthesis process. For example, automatically rotating the textures so their repeating patterns are aligned with the axes was studied by [Rod+19]. Powerful methods for artifact [Dek+15] and distortion [Li+19] removal could be applied as a pre-processing operation to the input textures before training the generative models, or as additional components to their loss function for improving tileability or homogeneity.

3.A Additional Details

Training details: We use PyTorch [Pas+19] as our learning framework, and Adam [KB15] as our optimization algorithm. For it, we use an initial learning rate of α = 0.0002, which is divided by 5, after iterations 30000 and 40000. The networks are trained for a total of 50000 iterations, using a batch size of 1.The momentum parameters are kept to the default values of β1=0.9 and β2=0.999, and ϵ=108. We additionally apply an L2 regularization of 10−5 to the optimization of both networks, which may help obtain better synthesized images [Kur+19]. This configuration is the same for the training of both the discriminator and the generator. In contrast to other work in unsupervised image generation [Zhu+17a], both networks are updated after each iteration. The models are trained and evaluated using a single NVIDIA GeForce GTX 1080Ti, leveraging half-precision training and automatic gradient scaling [Mic+18]. Evaluation is done in half precision for faster inference and reduced memory consumption. Using reduced numeric precision does not yield worse synthesized textures.

Network architecture: We follow closely the architectures defined in [Zho+18]. For the Generator G, we use replication padding as the padding method for the entire architecture, ReLU non-linearities, and Instance Normalization [UVL16] before each non-linearity. This network receives textures of 3×k×k pixels, which are first processed by an encoder E, which transforms it to a 256×k4×k4 latent vector x0. Following standard practice on residual convolutional networks [He+16], the first layer of E is comprised by 64 7 × 7 convolutional filters. The rest of the network, unless stated otherwise, uses standard 3 × 3 kernels. This latent vector is further processed by 5 residual blocks xiRi(xi1)+xi1, which maintain the latent vector dimensions. Then, a decoder G takes x5 and, using three convolutional layers with transposed convolutions [LSD15], followed by ReLU operations, it returns the estimated texture, with 2k×2k resolution. As in the encoder, the last layer of G contains 64 7 × 7 convolutional filters. There are as many decoders G as individual maps in the texture stack. When tiling the latent space x0, this doubles it spatial resolution, thus becoming a 256×k2×k2 latent vector. As such, once processed by the residual blocks and the decoders, the network outputs a 4k×4k texture stack.

For the discriminator D, we simply follow a PatchGAN architecture [Zhu+17a; Iso+17; Led+17; Zho+18], which outputs local estimations of the probability of the input image being real or generated. We use 4 blocks of layers, each comprised of, in order, a convolutional layer with 4 × 4 kernels, an Instance Normalization [UVL16] layer, and a Leaky ReLU non-linearity [MHN13]. We use a stride of 2 in the first and last layers to decrease the dimensions of the latent representations. The first block of layers contains 64 trainable kernels, which are doubled after each block. The last layer thus contains 512 trainable filters. In the last layer, we remove the normalization operation and substitute the non-linearity by a Sigmoid operation, to transform the output into the (0, 1) range. The discriminator receives 2k×2k resolution stacks, and outputs a 1×k2×k2 resolution estimation map.

3.B Additional Comparisons, Results and Ablations

In this section, we outline the implementation details used for our comparison with other tileable texture synthesis methods, which included Histogram-Preserving Blending by Deliot and Heitz [DH19], Graph-Cuts synthesis by Li et al. [Li+20], Repeated Pattern Detection by Rodriguez-Pardo et al. [Rod+19], Periodic Spatial GAN by Bergmann et al. [BJV17] and Self-Organizing Textures by Niklasson et al. [Nik+21]. We also provide further results in Figure 3.15 and an ablation study on the size of the model architectures in Figure 3.14.

Figure 3.14: An ablation study on the influence of the number of layers in each model in the quality of the generated outputs. We evaluate each network using the same input, marked using a blue box. We compare evaluate the influence of using discriminators with L ∈ {3, 4, 5} layers, marked as DL and generators with L ∈ {4, 5, 6} residual blocks, marked as GL. We show the configuration we use for every experiment in this chapter using a green box. On the bottom of each image, we show the training time for 50000 iterations. As shown, using a discriminator with 3 layers yields repetitive large-scale patterns, due to its reduced receptive field. Using DL = 5 yields marginal improvements with respect to DL = 4, with the cost of increased runtimes. In terms of GL, we find that there are no significant visual differences between using 5 or 6 residual blocks, but we favor the former due to its reduced training time.

Figure 3.14: An ablation study on the influence of the number of layers in each model in the quality of the generated outputs. We evaluate each network using the same input, marked using a blue box. We compare evaluate the influence of using discriminators with L ∈ {3, 4, 5} layers, marked as DL and generators with L ∈ {4, 5, 6} residual blocks, marked as GL. We show the configuration we use for every experiment in this chapter using a green box. On the bottom of each image, we show the training time for 50000 iterations. As shown, using a discriminator with 3 layers yields repetitive large-scale patterns, due to its reduced receptive field. Using DL = 5 yields marginal improvements with respect to DL = 4, with the cost of increased runtimes. In terms of GL, we find that there are no significant visual differences between using 5 or 6 residual blocks, but we favor the former due to its reduced training time.

Figure 3.15: More examples on the capabilities of SeamlessGAN on generating multiple tileable textures from a given input. On the left, we show the input sample T and four crops used for generating tileable maps. On the middle column, the outputs of the generative model are shown. On their right, the same textures are tiled 2 × 2 times for visualization purposes.

Figure 3.15: More examples on the capabilities of SeamlessGAN on generating multiple tileable textures from a given input. On the left, we show the input sample T and four crops used for generating tileable maps. On the middle column, the outputs of the generative model are shown. On their right, the same textures are tiled 2 × 2 times for visualization purposes.

Histogram-Preserving Blending We use the implementation provided by the authors. The only hyperparameter to be changed is the blending border size, which is set to the default value of 33% for all the textures.

Graph-Cuts synthesis To allow for comparison between single-map texture synthesis methods, we run this method to output only the albedo map, using the default values of λd = 1, a ratio of 0.64, a crop ratio of 0.6 and a margin size of 72 pixels. The images are not rescaled to any input or target size, instead, we use their original resolution.

Repeated Pattern Detection This paper proposed the use of some preprocessing methods for improving the regularity of the input images, which were not used for our comparison. We use their default hyperparameter values, including δ=0.65,λ=0.8, and the deep perceptual loss weighting proposed in [GEB15a].

Periodic Spatial GAN This method can handle both single-image and multiple image datasets. We are interested in the former, and, as such, each GAN is trained using only one texture sample. We train their model using Theano [Al-+16], and use their default configuration of a learning rate of 0.0002, a β1 = 0.5 as the momentum term for the optimization algorithm, a batch size of 25, and train their network for 10 epochs, each comprised of 25000 training steps. The network parameters and sizes are exactly as defined in the paper.

Self-Organizing Textures This method is computationally expensive and, as such, it requires that the input textures are rescaled to 128 × 128 pixels. We train their method using Tensorflow [Aba+16] for 5000 iterations. We use the default training configuration, with a starting learning rate of 0.002, which is then reduced by half at iterations 2000 and 4000. We do not modify the image parametrization design or loss function in any way.

Figure 3.16: Additional comparison with previous methods. From left to right, top to bottom: Input texture, the results of Histogram Blending, by Deloit et al. [DH19], the GraphCuts algorithm by Li et al. [Li+20], Texture Stationarization by Moritz et al. [Mor+17], Repeated Pattern Detection by Rodriguez-Pardo et al. [Rod+19], PSGAN by Bergmann et al. [BJV17], Self-Organizing Textures by Niklasson et al. [Nik+21]; and ours. Outputs are tiled a similar number of times (at least twice in each dimension) for better visualization. Our method generally captures better the overall structure of the texture, while providing seamless and semantically coherent borders, for enhancing tileability. From left to right, we sample c = 27, 36, 29, 1, 17, and 7 input crops before obtaining a tileable texture.

Figure 3.16: Additional comparison with previous methods. From left to right, top to bottom: Input texture, the results of Histogram Blending, by Deloit et al. [DH19], the GraphCuts algorithm by Li et al. [Li+20], Texture Stationarization by Moritz et al. [Mor+17], Repeated Pattern Detection by Rodriguez-Pardo et al. [Rod+19], PSGAN by Bergmann et al. [BJV17], Self-Organizing Textures by Niklasson et al. [Nik+21]; and ours. Outputs are tiled a similar number of times (at least twice in each dimension) for better visualization. Our method generally captures better the overall structure of the texture, while providing seamless and semantically coherent borders, for enhancing tileability. From left to right, we sample c = 27, 36, 29, 1, 17, and 7 input crops before obtaining a tileable texture.

BIBLIOGRAPHY

[Aba+16] Martın Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. “Tensorflow: A System for Large-Scale Machine Learning”. In: 12th Symposium on Operating Systems Design and Implementation. 2016, pp. 265–283 (cit. on p. 96).

[Aga+03] Sameer Agarwal, Ravi Ramamoorthi, Serge Belongie, and Henrik Wann Jensen. “Structured Importance Sampling of Environment Maps”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 605–612 (cit. on p. 185).

[AAL16] Miika Aittala, Timo Aila, and Jaakko Lehtinen. “Reflectance Modeling by Neural Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 35.4 (2016), pp. 1–13 (cit. on p. 124).

[AWL13] Miika Aittala, Tim Weyrich, and Jaakko Lehtinen. “Practical SVBRDF Capture in the Frequency Domain”. In: ACM Transactions on Graphics (TOG) 32.4 (2013), pp. 110–1 (cit. on pp. 17, 107, 128, 258).

[AWL15] Miika Aittala, Tim Weyrich, and Jaakko Lehtinen. “Two-shot SVBRDF Capture for Stationary Materials”. In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 1–13 (cit. on pp. 17, 34, 104, 125, 258).

[Akl+18] Adib Akl, Charles Yaacoub, Marc Donias, Jean-Pierre Da Costa, and Christian Germain. “A Survey of Exemplar-Based Texture Synthesis Methods”. In: Computer Vision and Image Understanding 172 (2018), pp. 12–24 (cit. on p. 73).

[Alc+19] Raul Alcain, Carlos Heras, Iñigo Salinas, Jorge Lopez-Moreno, and Carlos Aliaga. “Microscale Optical Capture System for Digital Fabric Recreation”. In: Proceedings of the 7th International Conference on Photonics, Optics and Laser Technology - Volume 1: PHOTOPTICS, INSTICC. SciTePress, 2019, pp. 114–119 (cit. on pp. 35, 39).

[Alm+18] Amjad Almahairi, Sai Rajeshwar, Alessandro Sordoni, Philip Bachman, and Aaron Courville. “Augmented Cyclegan: Learning Many-to-Many Mappings from Unpaired Data”. In: International Conference on Machine Learning (ICML). PMLR. 2018, pp. 195–204 (cit. on p. 85).

[Ami+20] Alexander Amini, Wilko Schwarting, Ava Soleimany, and Daniela Rus. “Deep Evidential Regression”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 14927–14937 (cit. on p. 126).

[AP08] Xiaobo An and Fabio Pellacini. “AppProp: All-pairs Appearance-Space Edit Propagation”. In: ACM Transactions on Graphics (TOG) 27.3 (2008), pp. 1–9 (cit. on p. 33).

[And+20] Pontus Andersson, Jim Nilsson, Tomas Akenine-Möller, Magnus Oskarsson, Kalle Åström, and Mark D Fairchild. “FLIP: A Difference Evaluator for Alternating Images.” In: Proc. ACM Comput. Graph. Interact. Tech. 3.2 (2020), pp. 15–1 (cit. on pp. 198, 199, 202).

[ASE17] Antreas Antoniou, Amos Storkey, and Harrison Edwards. “Data Augmentation Generative Adversarial Networks”. In: arXiv preprint arXiv:1711.04340 (2017) (cit. on p. 45).

[ACB19] Devansh Arpit, Vıctor Campos, and Yoshua Bengio. “How to Initialize Your Network? Robust Initialization for Weightnorm & Resnets”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on pp. 109, 119).

[ARV19] Yuki M Asano, Christian Rupprecht, and Andrea Vedaldi. “A Critical Analysis of Self-Supervision, or What We Can Learn from a Single Image”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 75).

[Att+22] Benjamin Attal, Jia-Bin Huang, Michael Zollhöfer, Johannes Kopf, and Changil Kim. “Learning Neural Light Fields with Ray-space Embedding”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 19819–19829 (cit. on pp. 187, 203, 212).

[Azi+19] Dejan Azinovic, Tzu-Mao Li, Anton Kaplanyan, and Matthias Nießner. “Inverse Path Tracing for Joint Material and Lighting Estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 2447–2456 (cit. on p. 203).

[Azi+23] Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia, Mohammad Norouzi, and David J Fleet. “Synthetic Data from Diffusion Models Improves ImageNet Classification”. In: arXiv preprint arXiv:2304.08466 (2023) (cit. on p. 212).

[BKH16] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. “Layer Normalization”. In: arXiv preprint arXiv:1607.06450 (2016) (cit. on pp. 109, 111, 119, 176).

[Baa+22] Hendrik Baatz, Jonathan Granskog, Marios Papas, Fabrice Rousselle, and Jan Novák. “NeRF-Tex: Neural Reflectance Field Textures”. In: Computer Graphics Forum. Vol. 41. 6. Wiley Online Library. 2022, pp. 287–301 (cit. on pp. 104, 187).

[BSK23] Steve Bako, Pradeep Sen, and Anton Kaplanyan. “Deep Appearance Prefiltering”. In: ACM Transactions on Graphics (TOG) 42.2 (2023), pp. 1–23 (cit. on p. 105).

[BYK21] Wentao Bao, Qi Yu, and Yu Kong. “Evidential Deep Learning for Open Set Action Recognition”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 13349–13358 (cit. on p. 130).

[Bar+09] Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman. “PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing”. In: ACM Transactions on Graphics (TOG) 28.3 (2009), p. 24 (cit. on p. 74).

[Bar+15] Connelly Barnes, Fang-Lue Zhang, Liming Lou, Xian Wu, and Shi-Min Hu. “Patchtable: Efficient Patch Queries for Large Datasets and Applications”. In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 1–10 (cit. on pp. 32, 74).

[Bar+21] Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. “Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021 (cit. on p. 186).

[Ben+21] Saguy Benaim, Ron Mokady, Amit Bermano, and Lior Wolf. “Structural Analogy from a Single Image Pair”. In: Computer Graphics Forum. Vol. 40. 1. Wiley Online Library. 2021, pp. 249–265 (cit. on pp. 31, 33, 46, 47, 56, 57, 60, 75).

[Bén+13] Pierre Bénard, Forrester Cole, Michael Kass, Igor Mordatch, James Hegarty, Martin Sebastian Senn, Kurt Fleischer, Davide Pesare, and Katherine Breeden. “Stylizing Animation by Example”. In: ACM Transactions on Graphics (TOG) 32.4 (2013), pp. 1–12 (cit. on p. 32).

[Ben+20] Gregory Benton, Marc Finzi, Pavel Izmailov, and Andrew Gordon Wilson. “Learning Invariances in Neural Networks”. In: arXiv preprint arXiv:2010.11882 (2020) (cit. on p. 43).

[BJV17] Urs Bergmann, Nikolay Jetchev, and Roland Vollgraf. “Learning Texture Manifolds with the Periodic Spatial GAN”. In: International Conference on Machine Learning (ICML). 2017, pp. 469–477 (cit. on pp. 72, 74–76, 89–91, 95, 99).

[Ber+23] Hugo Bertiche, Niloy J Mitra, Kuldeep Kulkarni, Chun-Hao Paul Huang, Tuanfeng Y Wang, Meysam Madadi, Sergio Escalera, and Duygu Ceylan. “Blowing in the Wind: CycleNet for Human Cinemagraphs from Still Images”. In: arXiv preprint arXiv:2303.08639 (2023) (cit. on pp. 17, 259).

[Bha+03] Kiran S. Bhat, Christopher D. Twigg, Jessica K. Hodgins, Pradeep K. Khosla, Zoran Popovic, and Steven M. Seitz. “Estimating Cloth Simulation Parameters from Video”. In: Symposium on Computer Animation. The Eurographics Association, 2003 (cit. on pp. 156, 158).

[Bi+18] Wenyan Bi, Peiran Jin, Hendrikje Nienborg, and Bei Xiao. “Estimating Mechanical Properties of Cloth from Videos Using Dense Motion Trajectories: Human Psychophysics and Machine Learning”. In: Journal of vision 18.5 (2018), pp. 12–12 (cit. on p. 159).

[Bie20] Lukas Biewald. Experiment Tracking with Weights and Biases. Software available from wandb.com. 2020 (cit. on pp. 108, 176, 194).

[Bit+20] Benedikt Bitterli, Chris Wyman, Matt Pharr, Peter Shirley, Aaron Lefohn, and Wojciech Jarosz. “Spatiotemporal Reservoir Resampling for Real-time Ray Tracing with Dynamic Direct Lighting”. In: ACM Transactions on Graphics (Proceedings of SIGGRAPH) 39.4 (2020) (cit. on p. 184).

[Blu+16] Adrian Blumer, Jan Novák, Ralf Habel, Derek Nowrouzezahrai, and Wojciech Jarosz. “Reduced Aggregate Scattering Operators for Path Tracing”. In: Computer Graphics Forum (Proceedings of Pacific Graphics) 35.7 (2016), pp. 461–473 (cit. on p. 203).

[Bom+21] Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. “On the Opportunities and Risks of Foundation Models”. In: (2021) (cit. on p. 159).

[Bos+20] Mark Boss, Varun Jampani, Kihwan Kim, Hendrik Lensch, and Jan Kautz. “Two-shot Spatially-Varying BRDF and Shape Estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 3982–3991 (cit. on p. 125).

[Bou+13] Katherine L Bouman, Bei Xiao, Peter Battaglia, and William T Freeman. “Estimating the Material Properties of Fabric from Video”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2013, pp. 1984–1991 (cit. on pp. 156, 159).

[BS12] Brent Burley and Walt Disney Animation Studios. “Physically-based Shading at Disney”. In: ACM SIGGRAPH. Vol. 2012. vol. 2012. 2012, pp. 1–7 (cit. on pp. 76, 127).

[Byl+22] Zoya Bylinskii, Laura Herman, Aaron Hertzmann, Stefanie Hutka, and Yile Zhang. “Towards Better User Studies in Computer Graphics and Vision”. In: arXiv preprint arXiv:2206.11461 (2022) (cit. on pp. 178, 182).

[CLA19] Carlos Castillo, Jorge López-Moreno, and Carlos Aliaga. “Recent Advances in Fabric Appearance Reproduction”. In: Computers & Graphics 84 (2019), pp. 103–121 (cit. on pp. 14, 17, 40, 256, 259).

[CHB21] Thomas Chambon, Eric Heitz, and Laurent Belcour. “Passing Multi-Channel Material Textures to a 3-Channel Loss”. In: ACM SIGGRAPH 2021 Talks. 2021, pp. 1–2 (cit. on pp. 64, 66, 67, 128).

[Cha+19] Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A. Efros. “Everybody Dance Now”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on p. 85).

[Che+20a] Chengqian Che, Fujun Luan, Shuang Zhao, Kavita Bala, and Ioannis Gkioulekas. “Towards Learning-based Inverse Subsurface Scattering”. In: 2020 IEEE International Conference on Computational Photography (ICCP). IEEE. 2020, pp. 1–12 (cit. on p. 203).

[Che+22] Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. “Tensorf: Tensorial Radiance Fields”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 333–350 (cit. on pp. 203, 212).

[Che+17a] Dongdong Chen, Jing Liao, Lu Yuan, Nenghai Yu, and Gang Hua. “Coherent Online Video Style Transfer”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 1105–1114 (cit. on p. 32).

[Che+17b] Dongdong Chen, Lu Yuan, Jing Liao, Nenghai Yu, and Gang Hua. “Stylebank: An Explicit Representation for Neural Image Style Transfer”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 1897–1906 (cit. on p. 32).

[Che+20b] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. “A Simple Framework for Contrastive Learning of Visual Representations”. In: International Conference on Machine Learning (ICML). PMLR. 2020, pp. 1597–1607 (cit. on pp. 37, 45, 159).

[CXH21] Xinlei Chen, Saining Xie, and Kaiming He. “An Empirical Study of Training Self-supervised Vision Transformers”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2021, pp. 9640–9649 (cit. on p. 159).

[Che+18] Yuanqin Chen, Qian Zhang, Yaping Wu, Bo Liu, Meiyun Wang, and Yusong Lin. “Fine-tuning ResNet for Breast Cancer Classification from mammography”. In: The International Conference on Healthcare Science and Engineering. Springer. 2018, pp. 83–96 (cit. on p. 159).

[CNN22] Zhe Chen, Shohei Nobuhara, and Ko Nishino. “Invertible Neural BRDF for Object Inverse Rendering”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 44.12 (2022), pp. 9380–9395 (cit. on pp. 186, 203).

[Cho+18] Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. “Stargan: Unified Generative Adversarial Networks for Multi-domain Image-to-image Translation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8789–8797 (cit. on p. 77).

[Cla+90] Timothy G Clapp, Hong Peng, Tushar K Ghosh, and Jeffrey W Eischen. “Indirect Measurement of the Moment-curvature Relationship for Fabrics”. In: Textile Research Journal 60.9 (1990), pp. 525–533 (cit. on pp. 156, 158).

[Cla+21] Petrik Clarberg, Wojciech Jarosz, Tomas Akenine-Möller, and Henrik Wann Jensen. “Wavelet Importance Sampling: Efficiently Evaluating Products of Complex Functions”. In: ACM Transactions on Graphics (Proceedings of SIGGRAPH) 24.3 (2021), pp. 3520–3528 (cit. on pp. 185, 189, 195, 199, 200).

[CTT17] David Clyde, Joseph Teran, and Rasmus Tamstorf. “Modeling and Data-driven Parameter Estimation for Woven Fabrics”. In: Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation. 2017, pp. 1–11 (cit. on p. 158).

[Coh+03] Michael F Cohen, Jonathan Shade, Stefan Hiller, and Oliver Deussen. “Wang Tiles for Image and Texture Generation”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 287–294 (cit. on p. 72).

[CT82] Robert L Cook and Kenneth E. Torrance. “A Reflectance Model for Computer Graphics”. In: ACM Transactions on Graphics (TOG) 1.1 (1982), pp. 7–24 (cit. on p. 48).

[Dan01] Kristin J Dana. “BRDF/BTF Measurement Device”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Vol. 2. IEEE. 2001, pp. 460–466 (cit. on p. 103).

[Dan+99] Kristin J Dana, Bram Van Ginneken, Shree K Nayar, and Jan J Koenderink. “Reflectance and Texture of Real-World Surfaces”. In: ACM Transactions On Graphics (TOG) 18.1 (1999), pp. 1–34 (cit. on p. 33).

[Dar+12] Soheil Darabi, Eli Shechtman, Connelly Barnes, Dan B Goldman, and Pradeep Sen. “Image Melding: Combining Inconsistent Images Using Patch-Based Synthesis”. In: ACM Transactions on Graphics (TOG) 31.4 (2012), pp. 1–10 (cit. on p. 74).

[Dav+15] Abe Davis, Katherine L Bouman, Justin G Chen, Michael Rubinstein, Fredo Durand, and William T Freeman. “Visual Vibrometry: Estimating Material Properties from Small Motion in Video”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2015, pp. 5335–5343 (cit. on p. 159).

[De 97] Jeremy S De Bonet. “Multiresolution Sampling Procedure for Analysis and Synthesis of Texture Images”. In: Proceedings of the 24th Annual cCnference on Computer Graphics and Interactive Techniques. 1997, pp. 361–368 (cit. on p. 74).

[Dek+15] Tali Dekel, Tomer Michaeli, Michal Irani, and William T. Freeman. “Revealing and Modifying Non-Local Variations in a Single Image”. In: ACM Transactions on Graphics (TOG) (2015) (cit. on p. 93).

[DH19] Thomas Deliot and Eric Heitz. “Procedural Stochastic Textures by Tiling and Blending”. In: GPU Zen 2 (2019) (cit. on pp. 72, 74, 89–91, 95, 99).

[Den+09] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. “Imagenet: A Large-scale Hierarchical Image Database”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2009, pp. 248–255 (cit. on pp. 32, 164, 166).

[Des+18] Valentin Deschaintre, Miika Aittala, Fredo Durand, George Drettakis, and Adrien Bousseau. “Single-image SVBRDF Capture with a Rendering-Aware Deep Network”. In: ACM Transactions on Graphics (TOG) 37.4 (2018) (cit. on pp. 17, 34, 39, 54, 77, 102, 123, 125, 127, 258).

[Des+19] Valentin Deschaintre, Miika Aittala, Frédo Durand, George Drettakis, and Adrien Bousseau. “Flexible SVBRDF Capture with a Multi-Image Deep Network”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 1–13 (cit. on pp. 34, 42, 72, 76, 102, 125, 141).

[DDB20] Valentin Deschaintre, George Drettakis, and Adrien Bousseau. “Guided Fine-Tuning for Large-Scale Material Transfer”. In: Computer Graphics Forum. Vol. 39. 4. Wiley Online Library. 2020, pp. 91–105 (cit. on pp. 35, 38–40, 47, 49, 50, 63, 72, 76, 105, 125, 142).

[DLG21] Valentin Deschaintre, Yiming Lin, and Abhijeet Ghosh. “Deep Polarization Imaging for 3D Shape and SVBRDF Acquisition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 15567–15576 (cit. on pp. 64, 67, 69, 127, 141).

[Dia+20] Foivos I Diakogiannis, François Waldner, Peter Caccetta, and Chen Wu. “ResUNet-a: A Deep Learning Framework for Semantic Segmentation of Remotely Sensed Data”. In: ISPRS Journal of Photogrammetry and Remote Sensing 162 (2020), pp. 94–114 (cit. on pp. 64, 66, 69, 108, 127, 141).

[Dia+15] Olga Diamanti, Connelly Barnes, Sylvain Paris, Eli Shechtman, and Olga Sorkine-Hornung. “Synthesis of Complex Image Appearance from Limited Exemplars”. In: ACM Transactions on Graphics (TOG) 34.2 (2015), pp. 1–14 (cit. on pp. 34, 104).

[Din+20] Keyan Ding, Kede Ma, Shiqi Wang, and Eero P Simoncelli. “Image Quality Assessment: Unifying Structure and Texture Similarity”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence (2020) (cit. on p. 160).

[DKB14] Laurent Dinh, David Krueger, and Yoshua Bengio. “Nice: Non-linear Independent Components Estimation”. In: arXiv preprint arXiv:1410.8516 (2014) (cit. on p. 191).

[DSB16] Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. “Density Estimation Using Real NVP”. In: arXiv preprint arXiv:1605.08803 (2016) (cit. on p. 191).

[Dod+22] Ana Dodik, Silvia Sellán, Theodore Kim, and Amanda Phillips. “Sex and Gender in the Computer Graphics Research Literature”. In: arXiv preprint arXiv:2206.00480 (2022) (cit. on pp. 178, 179, 182).

[DZP07] Weiming Dong, Ning Zhou, and Jean-Claude Paul. “Optimized Tile-Based Texture Synthesis”. In: Proceedings of Graphics Interface 2007. 2007, pp. 249–256 (cit. on p. 74).

[Don19] Yue Dong. “Deep Appearance Modeling: A Survey”. In: Visual Informatics 3.2 (2019), pp. 59–68 (cit. on pp. 34, 53).

[Don+10] Yue Dong, Jiaping Wang, Xin Tong, John Snyder, Yanxiang Lan, Moshe Ben-Ezra, and Baining Guo. “Manifold Bootstrapping for SVBRDF Capture”. In: ACM Transactions on Graphics (TOG) 29.4 (2010), pp. 1–10 (cit. on p. 34).

[DB16] Alexey Dosovitskiy and Thomas Brox. “Generating images with Perceptual Similarity Metrics based on Deep Networks”. In: Advances in Neural Information Processing Systems. 2016, pp. 658–666 (cit. on p. 74).

[DTD21] Emilien Dupont, Yee Whye Teh, and Arnaud Doucet. “Generative Models as Distributions of Functions”. In: arXiv preprint arXiv:2102.04776 (2021) (cit. on p. 75).

[Dur+19a] Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. “Neural Spline Flows”. In: Advances in Neural Information Processing Systems. Vol. 32. 2019 (cit. on pp. 185, 191, 192, 194, 196, 197).

[Dur+19b] Conor Durkan, Artur Bekasovs, Iain Murray, and Georgios Papamakarios. “Cubic-Spline Flows”. In: Workshop on Invertible Neural Nets and Normalizing Flows: ICML 2019. 2019 (cit. on p. 191).

[Dye+18] Joanne Dyer, Diego Tamburini, Elisabeth R O’Connell, and Anna Harrison. “A Multispectral Imaging Approach Integrated into the Study of Late Antique textiles from Egypt”. In: PLoS One 13.10 (2018), e0204699 (cit. on p. 213).

[EF01] Alexei A Efros and William T Freeman. “Image Quilting for Texture Synthesis and Transfer”. In: Proceedings of the 28th annual conference on Computer Graphics and Interactive Techniques. 2001, pp. 341–346 (cit. on pp. 72, 74).

[EL99] Alexei A Efros and Thomas K Leung. “Texture Synthesis by Non-parametric Sampling”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Vol. 2. IEEE. 1999, pp. 1033–1038 (cit. on p. 74).

[EM17] Michael Elad and Peyman Milanfar. “Style Transfer via Texture Synthesis”. In: IEEE Transactions on Image Processing 26.5 (2017), pp. 2338–2351 (cit. on pp. 34, 104).

[EUD18] Stefan Elfwing, Eiji Uchibe, and Kenji Doya. “Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning”. In: Neural Networks 107 (2018), pp. 3–11 (cit. on pp. 141, 144).

[End+16] Yuki Endo, Satoshi Iizuka, Yoshihiro Kanamori, and Jun Mitani. “Deepprop: Extracting Deep Features from a Single Image for Edit Propagation”. In: Computer Graphics Forum. Vol. 35. 2. Wiley Online Library. 2016, pp. 189–201 (cit. on p. 33).

[Fan+22] Jiahui Fan, Beibei Wang, Miloš Hašan, Jian Yang, and Ling-Qi Yan. “Neural Layered BRDFs”. In: Proceedings of SIGGRAPH 2022. 2022 (cit. on pp. 116, 187).

[Fen+22] Xudong Feng, Wenchao Huang, Weiwei Xu, and Huamin Wang. “Learning-Based Bending Stiffness Parameter Estimation by a Drape Tester”. In: ACM Transactions on Graphics (TOG) 41.6 (2022), pp. 1–16 (cit. on p. 159).

[Fil+18] J. Filip, M. Kolafová, M. Havlıček, R. Vávra, M. Haindl, and Rushmeier H. “Evaluating Physical and Rendered Material Appearance”. In: The Visual Computer (Computer Graphics International 2018) (2018) (cit. on p. 109).

[FH08] Jiřı Filip and Michal Haindl. “Bidirectional Texture Function Modeling: A State of the Art Survey”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 31.11 (2008), pp. 1921–1940 (cit. on p. 103).

[FR22] Michael Fischer and Tobias Ritschel. “Plateau-Reduced Differentiable Path Tracing”. In: arXiv preprint arXiv:2211.17263 (2022) (cit. on p. 202).

[Fiš+16] Jakub Fišer, Ondřej Jamriška, Michal Lukáč, Eli Shechtman, Paul Asente, Jingwan Lu, and Daniel Sy`kora. “StyLit: Illumination-Guided Example-Based Stylization of 3D Renderings”. In: ACM Transactions on Graphics (TOG) 35.4 (2016), pp. 1–11 (cit. on p. 33).

[Fri+22] Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. “Plenoxels: Radiance Fields Without Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 5501–5510 (cit. on pp. 187, 212).

[FAW19] Anna Frühstück, Ibraheem Alhashim, and Peter Wonka. “TileGAN: Synthesis of Large-Scale Non-Homogeneous Textures”. In: ACM Transactions on Graphics (TOG) 38.4 (Apr. 2019) (cit. on pp. 34, 72, 74, 79–81, 104).

[Fu+20] Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yulan Guo, Yinghui Gao, and Biao Li. “Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs”. In: 31st British Machine Vision Conference 2020, BMVC 2020, BMVA Press, 2020 (cit. on pp. 167, 168).

[GG16] Yarin Gal and Zoubin Ghahramani. “Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning”. In: International Conference on Machine Learning (ICML). PMLR. 2016, pp. 1050–1059 (cit. on pp. 123, 126, 130, 141).

[Gal+12] Bruno Galerne, Ares Lagae, Sylvain Lefebvre, and George Drettakis. “Gabor Noise by Example”. In: ACM Transactions on Graphics (TOG) 31.4 (2012), pp. 1–9 (cit. on p. 72).

[Gao+21] Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. “Dynamic View Synthesis From Dynamic Monocular Video”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 5712–5721 (cit. on p. 186).

[Gao+19] Duan Gao, Xiao Li, Yue Dong, Pieter Peers, Kun Xu, and Xin Tong. “Deep Inverse Rendering for High-Resolution SVBRDF Estimation from an Arbitrary Number of Images”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–15 (cit. on pp. 49, 59, 61, 125, 137, 138, 152, 153).

[GMX22] Duan Gao, Haoyuan Mu, and Kun Xu. “Neural Global Illumination: Interactive In-direct Illumination Prediction under Dynamic Area Lights”. In: IEEE Transactions on Visualization and Computer Graphics (2022) (cit. on pp. 105, 186, 203).

[Gar+14] Elena Garces, Aseem Agarwala, Diego Gutierrez, and Aaron Hertzmann. “A Similarity Measure for Illustration Style”. In: ACM Transactions on Graphics (TOG) 33.4 (2014), pp. 1–9 (cit. on pp. 157, 160, 170, 178, 180).

[Gar+23] Elena Garces, Victor Arellano, Carlos Rodriguez-Pardo, David Pascual-Hernandez, Sergio Suja, and Jorge Lopez-Moreno. “Towards Material Digitization with a Dual-scale Optical System”. In: ACM Transactions on Graphics (TOG) (2023) (cit. on pp. 15, 16, 24, 102, 109, 269).

[Gar+22] Elena Garces, Carlos Rodriguez-Pardo, Dan Casas, and Jorge Lopez-Moreno. “A Survey on Intrinsic Images: Delving Deep into Lambert and Beyond”. In: International Journal of Computer Vision (2022) (cit. on pp. 3, 4, 6, 7, 12, 22, 23, 127, 128, 141, 268).

[GES22] James Gardner, Bernhard Egger, and William Alfred Peter Smith. “Rotation-Equivariant Conditional Spherical Neural Fields for Learning a Natural Illumination Prior”. In: Advances in Neural Information Processing Systems. 2022 (cit. on pp. 19, 185, 186, 193, 196, 199, 202, 203, 260).

[Gar+19] Marc-André Gardner, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Christian Gagné, and Jean-François Lalonde. “Deep Parametric Indoor Lighting Estimation”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019, pp. 7175–7183 (cit. on p. 203).

[GEB15a] Leon Gatys, Alexander S Ecker, and Matthias Bethge. “Texture Synthesis using Convolutional Neural Networks”. In: Advances in Neural Information Processing Systems. 2015, pp. 262–270 (cit. on pp. 74, 95, 128).

[GEB15b] Leon A Gatys, Alexander S Ecker, and Matthias Bethge. “A Neural Algorithm of Artistic Style”. In: arXiv preprint arXiv:1508.06576 (2015) (cit. on pp. 32, 47, 48, 124, 128).

[GEB16a] Leon A Gatys, Alexander S Ecker, and Matthias Bethge. “Image Style Transfer using Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2016, pp. 2414–2423 (cit. on pp. 74, 160).

[Gat+17] Leon A Gatys, Alexander S Ecker, Matthias Bethge, Aaron Hertzmann, and Eli Shechtman. “Controlling Perceptual Factors in Neural Style Transfer”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 3985–3993 (cit. on pp. 32, 74).

[GEB16b] Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. “Image Style Transfer Using Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 2016-Decem. June 2016, pp. 2414–2423 (cit. on p. 79).

[Gau+22] Alban Gauthier, Robin Faury, Jérémy Levallois, Theo Thonat, Jean-Marc Thiery, and Tamy Boubekeur. “MIPNet: Neural Normal-to-Anisotropic-Roughness MIP mapping”. In: ACM Transactions on Graphics (TOG) 41.6 (2022), pp. 1–12 (cit. on p. 105).

[Gaw+21] Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al. “A Survey of Uncertainty in Deep Neural Networks”. In: arXiv preprint arXiv:2107.03342 (2021) (cit. on p. 126).

[Geo+18] Iliyan Georgiev, Thiago Ize, Mike Farnsworth, Ramón Montoya-Vozmediano, Alan King, Brecht Van Lommel, Angel Jimenez, Oscar Anson, Shinji Ogaki, Eric Johnston, et al. “Arnold: A Brute-force Production Path Tracer”. In: ACM Transactions on Graphics (TOG) 37.3 (2018), pp. 1–12 (cit. on p. 48).

[Gil+14] Guillaume Gilet, Basile Sauvage, Kenneth Vanhoey, Jean-Michel Dischler, and Djamchid Ghazanfarpour. “Local Random-phase Noise for Procedural Texturing”. In: ACM Transactions on Graphics (TOG) 33.6 (2014), pp. 1–11 (cit. on p. 72).

[Goo+14a] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. “Generative Adversarial Nets”. In: Advances in Neural Information Processing Systems. 2014, pp. 2672–2680 (cit. on p. 39).

[Goo+14b] Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. “Generative Adversarial Nets”. In: Advances in Neural Information Processing Systems. Vol. 3. January. Neural information processing systems foundation, June 2014, pp. 2672–2680 (cit. on p. 79).

[Gri+03] Eitan Grinspun, Anil N Hirani, Mathieu Desbrun, and Peter Schröder. “Discrete Shells”. In: Proceedings of the 2003 ACM SIGGRAPH/Eurographics symposium on Computer animation. Citeseer. 2003, pp. 62–67 (cit. on pp. 161, 162).

[Gro+20] Aditya Grover, Christopher Chute, Rui Shu, Zhangjie Cao, and Stefano Ermon. “Align-flow: Cycle Consistent Learning from Multiple Domains via Normalizing Flows”. In: Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 34. 04. 2020, pp. 4028–4035 (cit. on p. 186).

[Gu+18] Shuyang Gu, Congliang Chen, Jing Liao, and Lu Yuan. “Arbitrary Style Transfer with Deep Feature Reshuffle”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8222–8231 (cit. on p. 32).

[Gua+16] Darya Guarnera, Giuseppe Claudio Guarnera, Abhijeet Ghosh, Cornelia Denk, and Mashhuda Glencross. “BRDF Representation and Acquisition”. In: Computer Graphics Forum. Vol. 35. 2. Wiley Online Library. 2016, pp. 625–650 (cit. on p. 34).

[Gua+19] Giuseppe Claudio Guarnera, Dar’ya Guarnera, Gregory J Ward, Mashhuda Glencross, and Ian Hall. “BxDF Material Acquisition, Representation, and Rendering for VR and Design”. In: SIGGRAPH Asia 2019 Courses. 2019, pp. 1–21 (cit. on p. 3).

[Gue+20] P Guehl, R Allègre, J-M Dischler, B Benes, and E Galin. “Semi-Procedural Textures Using Point Process Texture Basis Functions”. In: Computer Graphics Forum. Vol. 39. 4. Wiley Online Library. 2020, pp. 159–171 (cit. on pp. 34, 72, 92, 104).

[Gui+17] Geoffrey Guingo, Basile Sauvage, Jean-Michel Dischler, and Marie-Paule Cani. “Bilayer Textures: A Model for Synthesis and Deformation of Composite Textures”. In: Computer Graphics Forum. Vol. 36. 4. 2017, pp. 111–122 (cit. on p. 72).

[Guo+21] Jie Guo, Shuichang Lai, Chengzhi Tao, Yuelong Cai, Lei Wang, Yanwen Guo, and Ling-Qi Yan. “Highlight-Aware Two-Stream Network for Single-Image SVBRDF Acquisition”. In: ACM Transactions on Graphics (TOG) 40.4 (2021), pp. 1–14 (cit. on pp. 123, 125).

[Guo+20a] Yu Guo, Miloš Hašan, Lingqi Yan, and Shuang Zhao. “A Bayesian Inference Framework for Procedural Material Parameter Estimation”. In: Computer Graphics Forum. Vol. 39. 7. Wiley Online Library. 2020, pp. 255–266 (cit. on pp. 125, 126).

[Guo+20b] Yu Guo, Cameron Smith, Miloš Hašan, Kalyan Sunkavalli, and Shuang Zhao. “MaterialGAN: Reflectance Capture using a Generative SVBRDF Model”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), 254:1–254:13 (cit. on pp. 17, 35, 42, 72, 76, 102, 123, 125, 258).

[Guo+19] Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris. “Spottune: Transfer Learning through Adaptive Fine-tuning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4805–4814 (cit. on p. 159).

[HDL16] David Ha, Andrew Dai, and Quoc V Le. “Hypernetworks”. In: arXiv preprint arXiv:1609.09106 (2016) (cit. on p. 75).

[HMV09] Simon Haegler, Pascal Müller, and Luc Van Gool. “Procedural Modeling for Digital Cultural Heritage”. In: EURASIP Journal on Image and Video Processing 2009 (2009), pp. 1–1 (cit. on p. 213).

[Hai12] Vávra R. Haindl M. Filip J. “Digital Material Appearance: the Curse of Tera-Bytes”. In: ERCIM News 90 (2012), pp. 49–50 (cit. on pp. 108, 109, 111).

[Ham+21] Hendrik Hameeuw, Godelieve Watteeuw, Bruno Vandermeulen, and Marc Proesmans. “The Painted Panels of the Early Sixteenth Century Mechelen Enclosed Gardens Art Technical Examination with the Photometric Stereo: White Light and Multispectral Microdomes”. In: Papers Presented at the Twentieth Symposium for the Study of Underdrawing and Technology in Painting held in Mechelen and Leuven, 11-13 January 2017. Vol. 20. Peeters Publishers; Leuven. 2021, pp. 165–175 (cit. on p. 213).

[Han+22a] Jiyeon Han, Hwanil Choi, Yunjey Choi, Junho Kim, Jung-Woo Ha, and Jaesik Choi. “Rarity Score: A New Metric to Evaluate the Uncommonness of Synthesized Images”. In: arXiv preprint arXiv:2206.08549 (2022) (cit. on p. 128).

[Han+22b] Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. “A Survey on Vision Transformer”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence (2022) (cit. on p. 164).

[Haq+23] Ayaan Haque, Matthew Tancik, Alexei Efros, Aleksander Holynski, and Angjoo Kanazawa. “Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions”. In: arXiv preprint arXiv:2303.12789 (2023) (cit. on p. 116).

[HHM22] Jon Hasselgren, Nikolai Hofmann, and Jacob Munkberg. “Shape, Light & Material Decomposition from Images Using Monte Carlo Rendering and Denoising”. In: arXiv preprint arXiv:2206.03380 (2022) (cit. on p. 203).

[HFM10] Vlastimil Havran, Jirı Filip, and Karol Myszkowski. “Bidirectional Texture Function Compression Based on Multi-level Vector Quantization”. In: Computer Graphics Forum. Vol. 29. 1. Wiley Online Library. 2010, pp. 175–190 (cit. on p. 103).

[He+22] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. “Masked Autoencoders are Scalable Vision Learners”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 16000–16009 (cit. on p. 159).

[He+16] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. “Deep Residual Learning for Image Recognition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 2016-Decem. 2016, pp. 770–778 (cit. on pp. 64, 66, 69, 78, 94, 108, 109, 127, 141, 164, 176).

[He+18] Mingming He, Dongdong Chen, Jing Liao, Pedro V Sander, and Lu Yuan. “Deep Exemplar-Based Colorization”. In: ACM Transactions on Graphics (TOG) 37.4 (2018), pp. 1–16 (cit. on p. 33).

[He+19] Mingming He, Jing Liao, Dongdong Chen, Lu Yuan, and Pedro V Sander. “Progressive Color Transfer with Dense Semantic Correspondences”. In: ACM Transactions on Graphics (TOG) 38.2 (2019), pp. 1–18 (cit. on p. 33).

[HB95] David J Heeger and James R Bergen. “Pyramid-based Texture Analysis/Synthesis”. In: Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques. 1995, pp. 229–238 (cit. on p. 74).

[Hei20] Eric Heitz. “Can’t Invert the CDF? The Triangle-Cut Parameterization of the Region under the Curve”. In: Computer Graphics Forum 39.4 (2020), pp. 121–132 (cit. on p. 185).

[HN18] Eric Heitz and Fabrice Neyret. “High-Performance By-Example Noise Using a Histogram-Preserving Blending Operator”. In: Proceedings of the ACM on Computer Graphics and Interactive Techniques 1.2 (2018), pp. 1–25 (cit. on p. 72).

[Hei+21] Eric Heitz, Kenneth Vanhoey, Thomas Chambon, and Laurent Belcour. “A Sliced Wasserstein Loss for Neural Texture Synthesis”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 9412–9420 (cit. on pp. 75, 93).

[Hel+20] Leonhard Helminger, Abdelaziz Djelouah, Markus Gross, and Christopher Schroers. “Lossy Image Compression with Normalizing Flows”. In: arXiv preprint arXiv:2008.10486 (2020) (cit. on p. 186).

[HG16] Dan Hendrycks and Kevin Gimpel. “Gaussian Error Linear Units (GELUs)”. In: arXiv preprint arXiv:1606.08415 (2016) (cit. on pp. 109, 119).

[Hen+21] Philipp Henzler, Valentin Deschaintre, Niloy J Mitra, and Tobias Ritschel. “Generative Modelling of BRDF Textures from Flash Images”. In: ACM Transactions on Graphics (Proc. SIGGRAPH Asia) 40.6 (2021) (cit. on pp. 17, 76, 102, 123, 124, 137, 138, 143, 151–153, 258).

[Her+20] Amir Hertz, Rana Hanocka, Raja Giryes, and Daniel Cohen-Or. “Deep Geometric Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 39.4 (2020), pp. 108–1 (cit. on p. 72).

[Her+01] Aaron Hertzmann, Charles E Jacobs, Nuria Oliver, Brian Curless, and David H Salesin. “Image Analogies”. In: Proceedings of the 28th annual conference on Computer graphics and interactive techniques. 2001, pp. 327–340 (cit. on pp. 31, 32, 46, 56, 57, 60).

[HS05] Aaron Hertzmann and Steven M Seitz. “Example-based Photometric Stereo: Shape Reconstruction with General, Varying BRDFs”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 27.8 (2005), pp. 1254–1264 (cit. on p. 34).

[HS03] Aaron Hertzmann and Steven M Seitz. “Shape and Materials by example: A Photometric Stereo Approach”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 1. IEEE. 2003, pp. I–I (cit. on p. 34).

[Heu+17] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. “Gans Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”. In: arXiv preprint arXiv:1706.08500 (2017) (cit. on p. 85).

[Hil+15] Stephen Hill, Stephen McAuley, Brent Burley, et al. “Physically Based Shading in Theory and Practice”. In: ACM SIGGRAPH 2015 Courses. SIGGRAPH ’15. Los Angeles, California: Association for Computing Machinery, 2015 (cit. on p. 11).

[HS06] Geoffrey E Hinton and Ruslan R Salakhutdinov. “Reducing the Dimensionality of Data with Neural Networks”. In: science 313.5786 (2006), pp. 504–507 (cit. on p. 103).

[Hin+21] Tobias Hinz, Matthew Fisher, Oliver Wang, and Stefan Wermter. “Improved Techniques for Training Single-Image GANs”. In: Proceedings - 2021 IEEE Winter Conference on Applications of Computer Vision, WACV 2021. Jan. 2021, pp. 1300–1309 (cit. on p. 75).

[HJA20] Jonathan Ho, Ajay Jain, and Pieter Abbeel. “Denoising Diffusion Probabilistic Models”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 6840–6851 (cit. on pp. 127, 209).

[HS13] Shahera Hossain and Seiichi Serikawa. “Texture Databases–A Comprehensive Survey”. In: Pattern Recognition Letters 34.15 (2013), pp. 2007–2022 (cit. on p. 109).

[Hu+20] Bingyang Hu, Jie Guo, Yanjun Chen, Mengtian Li, and Yanwen Guo. “DeepBRDF: A Deep Representation for Manipulating Measured BRDF”. In: Computer Graphics Forum. Vol. 39. 2. Wiley Online Library. 2020, pp. 157–166 (cit. on pp. 105, 203).

[Hu+22a] Dongting Hu, Liuhua Peng, Tingjin Chu, Xiaoxing Zhang, Yinian Mao, Howard Bondell, and Mingming Gong. “Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2022 (cit. on pp. 126, 130).

[HSS18] Jie Hu, Li Shen, and Gang Sun. “Squeeze-and-excitation Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 7132–7141 (cit. on p. 164).

[HXP20] Wei Hu, Lechao Xiao, and Jeffrey Pennington. “Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks”. In: arXiv preprint arXiv:2001.05992 (2020) (cit. on p. 141).

[HDR19] Yiwei Hu, Julie Dorsey, and Holly Rushmeier. “A Novel Framework for Inverse Procedural Texture Modeling”. In: ACM Transactions on Graphics (TOG) 38.6 (2019), pp. 1–14 (cit. on p. 72).

[Hu+22b] Yiwei Hu, Miloš Hašan, Paul Guerrero, Holly Rushmeier, and Valentin Deschaintre. “Controlling Material Appearance by Examples”. In: Computer Graphics Forum. Vol. 41. 4. Wiley Online Library. 2022, pp. 117–128 (cit. on p. 105).

[Hu+22c] Yiwei Hu, Chengan He, Valentin Deschaintre, Julie Dorsey, and Holly Rushmeier. “An Inverse Procedural Modeling Pipeline for SVBRDF Maps”. In: ACM Transactions on Graphics (TOG) 41.2 (2022), pp. 1–17 (cit. on p. 125).

[Hu+19] Yuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun, Nathan Carr, Jonathan Ragan-Kelley, and Fredo Durand. “Diff Taichi: Differentiable Programming for Physical Simulation”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 158).

[Hua+18a] Chin-Wei Huang, David Krueger, Alexandre Lacoste, and Aaron Courville. “Neural Autoregressive Flows”. In: International Conference on Machine Learning (ICML). PMLR. 2018, pp. 2078–2087 (cit. on p. 191).

[HB17] Xun Huang and Serge Belongie. “Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 1501–1510 (cit. on p. 32).

[Hua+18b] Xun Huang, Ming-Yu Liu, Serge Belongie, and Jan Kautz. “Multimodal Unsupervised Image-to-image Translation”. In: Proceedings of the European Conference on Computer Vision (ECCV). Sept. 2018 (cit. on p. 85).

[Hua+19] Zhiyuan Huang, Mansur Arief, Henry Lam, and Ding Zhao. “Evaluation Uncertainty in Data-Driven Self-Driving Testing”. In: 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE. 2019, pp. 1902–1907 (cit. on p. 126).

[HEW17] Markus Huber, Bernhard Eberhardt, and Daniel Weiskopf. “Cloth Animation Retrieval Using a Motion-shape Signature”. In: IEEE Computer Graphics and Applications 37.6 (2017), pp. 52–64 (cit. on p. 159).

[HZ21] Drew A Hudson and C Lawrence Zitnick. “Generative Adversarial Transformers”. In: arXiv preprint arXiv:2103.01209 (2021) (cit. on p. 93).

[Hwa+22] Inseung Hwang, Daniel S Jeon, Adolfo Muñoz, Diego Gutierrez, Xin Tong, and Min H Kim. “Sparse Ellipsometry: Portable Acquisition of Polarimetric SVBRDF and Shape with Unstructured Flash Photography”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), pp. 1–14 (cit. on p. 129).

[Ike81] Katsushi Ikeuchi. “Determining Surface Orientations of Specular Surfaces by Using the Photometric Stereo Method”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 6 (1981), pp. 661–669 (cit. on pp. 34, 40, 50, 84, 112, 151).

[IS15] Sergey Ioffe and Christian Szegedy. “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”. In: International Conference on Machine Learning (ICML). PMLR. 2015, pp. 448–456 (cit. on pp. 54, 64, 194).

[Iso+17] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. “Image-to-Image Translation with Conditional Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 5967–5976 (cit. on pp. 38, 39, 54, 56, 78, 79, 94, 107, 128, 132, 193).

[Jak+22] Wenzel Jakob, Sébastien Speierer, Nicolas Roussel, Merlin Nimier-David, Delio Vicini, Tizian Zeltner, Baptiste Nicolet, Miguel Crespo, Vincent Leroy, and Ziyi Zhang. Mitsuba 3 Renderer. Version 3.1.1. https://mitsuba-renderer.org. 2022 (cit. on pp. 189, 202).

[Jam+19] Ondřej Jamriška, Šárka Sochorová, Ondřej Texler, Michal Lukáč, Jakub Fišer, Jingwan Lu, Eli Shechtman, and Daniel Sy`kora. “Stylizing Video by Example”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–11 (cit. on p. 32).

[Jan+17] Michael Janner, Jiajun Wu, Tejas D Kulkarni, Ilker Yildirim, and Josh Tenenbaum. “Self-Supervised Intrinsic Image Decomposition”. In: Advances in Neural Information Processing Systems. 2017, pp. 5936–5946 (cit. on pp. 64, 67, 69, 78).

[JBH19] Miguel Jaques, Michael Burke, and Timothy Hospedales. “Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 158).

[JCJ09] Wojciech Jarosz, Nathan A. Carr, and Henrik Wann Jensen. “Importance Sampling Spherical Harmonics”. In: Computer Graphics Forum (Proceedings of Eurographics) 28.2 (2009), pp. 577–586 (cit. on p. 185).

[JBV16] Nikolay Jetchev, Urs Bergmann, and Roland Vollgraf. “Texture Synthesis with Spatial Generative Adversarial Networks”. In: arXiv preprint arXiv:1611.08207 (2016) (cit. on pp. 72, 74).

[Jia+21] Liming Jiang, Bo Dai, Wayne Wu, and Chen Change Loy. “Focal Frequency Loss for Image Reconstruction and Synthesis”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 13919–13929 (cit. on pp. 107, 128).

[Jin+22] Wenhua Jin, Beibei Wang, Miloš Hašan, Yu Guo, Steve Marschner, and Ling-Qi Yan. “Woven Fabric Capture from a Single Photo”. In: Proceedings of SIGGRAPH Asia 2022. 2022 (cit. on p. 125).

[Jin+19] Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, and Mingli Song. “Neural Style Transfer: A review”. In: IEEE Transactions on Visualization and Computer Graphics (2019) (cit. on p. 32).

[JAF16] Justin Johnson, Alexandre Alahi, and Li Fei-Fei. “Perceptual Losses for Real-time Style Transfer and Super-resolution”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2016, pp. 694–711 (cit. on pp. 32, 74).

[Jos+22] Laurent Valentin Jospin, Hamid Laga, Farid Boussaid, Wray Buntine, and Mohammed Bennamoun. “Hands-on Bayesian Neural Networks—A Tutorial for Deep Learning Users”. In: IEEE Computational Intelligence Magazine 17.2 (2022), pp. 29–48 (cit. on p. 126).

[JC20] Eunjung Ju and Myung Geol Choi. “Estimating Cloth Simulation Parameters From a Static Drape Using Neural Networks”. In: IEEE Access 8 (2020), pp. 195113–195121 (cit. on pp. 156, 159).

[KBH03] Florian Kainz, Rod Bogart, and Drew Hess. “The OpenEXR Image File Format”. In: ACM SIGGRAPH Technical Sketches (2003) (cit. on p. 185).

[Kaj86] James T Kajiya. “The Rendering Equation”. In: Proceedings of the 13th annual conference on Computer graphics and interactive techniques. 1986, pp. 143–150 (cit. on pp. 4, 8).

[Kam+16] Christos Kampouris, Stefanos Zafeiriou, Abhijeet Ghosh, and Sotiris Malassiotis. “Fine-Grained Material Classification Using Micro-Geometry and Reflectance”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2016, pp. 778–792 (cit. on p. 40).

[KG13] Brian Karis and Epic Games. “Real Shading in Unreal Engine 4”. In: Proc. Physically Based Shading Theory Practice 4.3 (2013), p. 1 (cit. on p. 127).

[Kar+18] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. “Progressive Growing of GANs for Improved Quality, Stability, and Variation”. In: International Conference on Learning Representations (ICLR). 2018 (cit. on pp. 74, 80).

[Kar+20a] Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. “Training Generative Adversarial Networks with Limited Data”. In: arXiv preprint arXiv:2006.06676 (2020) (cit. on pp. 36, 109).

[Kar+21] Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. “Alias-free Generative Adversarial Networks”. In: Advances in Neural Information Processing Systems. Vol. 34. 2021, pp. 852–863 (cit. on p. 186).

[KLA19] Tero Karras, Samuli Laine, and Timo Aila. “A Style-Based Generator Architecture for Generative Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4401–4410 (cit. on p. 74).

[Kar+20b] Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. “Analyzing and Improving the Image Quality of StyleGAN”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2020 (cit. on pp. 74, 85).

[Kas+15] Alexandre Kaspar, Boris Neubert, Dani Lischinski, Mark Pauly, and Johannes Kopf. “Self Tuning Texture Optimization”. In: Computer Graphics Forum. Vol. 34. 2. 2015, pp. 349–359 (cit. on p. 74).

[KZP19] Sergey Kastryulin, Dzhamil Zakirov, and Denis Prokopenko. PyTorch Image Quality: Metrics and Measure for Image Quality Assessment. Open-source software available at https://github.com/photosynthesis-team/piq. 2019 (cit. on pp. 182, 198).

[Kau17] Eric Kauderer-Abrams. “Quantifying Translation-invariance in Convolutional Neural Networks”. In: arXiv preprint arXiv:1801.01450 (2017) (cit. on p. 37).

[Kau+00] Jan Kautz, Pere-Pau Vázquez, Wolfgang Heidrich, and Hans-Peter Seidel. “Unified Approach to Prefiltered Environment Maps”. In: Proceedings of the Eurographics Workshop on Rendering Techniques 2000. 2000, pp. 185–196 (cit. on p. 185).

[Kaw80] Sueo Kawabata. “The Standardization and Analysis of Hand Evaluation”. In: The Textile Machinery Society of Japan (1980) (cit. on pp. 156, 158).

[KG17] Alex Kendall and Yarin Gal. “What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?” In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 126).

[Khe+21] Ilyes Khemakhem, Ricardo Monti, Robert Leech, and Aapo Hyvarinen. “Causal Autoregressive Flows”. In: International Conference on Artificial Intelligence and Statistics. PMLR. 2021, pp. 3520–3528 (cit. on p. 191).

[KB15] Diederik P. Kingma and Jimmy Ba. “Adam: A Method for Stochastic Optimization”. In: International Conference on Learning Representations (ICLR). 2015 (cit. on pp. 39, 55, 67, 83, 94, 110, 142, 194).

[Kin+16] Durk P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. “Improved Variational Inference with Inverse Autoregressive Flow”. In: Advances in Neural Information Processing Systems 29 (2016) (cit. on p. 191).

[KD18] Durk P. Kingma and Prafulla Dhariwal. “Glow: Generative Flow with Invertible 1x1 Convolutions”. In: Advances in Neural Information Processing Systems. Vol. 31. 2018 (cit. on pp. 186, 191).

[Kir+23] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. “Segment Anything”. In: arXiv preprint arXiv:2304.02643 (2023) (cit. on p. 212).

[KPB20] Ivan Kobyzev, Simon JD Prince, and Marcus A Brubaker. “Normalizing Flows: An Introduction and Review of Current Methods”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 43.11 (2020), pp. 3964–3979 (cit. on p. 191).

[Kol+20] Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. “Big Transfer (bit): General Visual Representation Learning”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 491–507 (cit. on p. 159).

[KSL19] Simon Kornblith, Jonathon Shlens, and Quoc V Le. “Do Better Imagenet Models Transfer Better?” In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 2661–2671 (cit. on p. 159).

[Kou+03] Melissa L Koudelka, Sebastian Magda, Peter N Belhumeur, and David J Kriegman. “Acquisition, Compression, and Synthesis of Bidirectional Texture Functions”. In: 3rd International Workshop on Texture Analysis and Synthesis (Texture 2003). 2003, pp. 59–64 (cit. on p. 103).

[Kry+21] Michael C Krygier, Tyler LaBonte, Carianne Martinez, Chance Norris, Krish Sharma, Lincoln N Collins, Partha P Mukherjee, and Scott A Roberts. “Quantifying the Unknown Impact of Segmentation Uncertainty on Image-Based Simulations”. In: Nature Communications 12.1 (2021), pp. 1–11 (cit. on p. 126).

[KLG20] Sandra Kuijpers, Christiane Luible-Bär, and R. Hugh Gong. “The Measurement of Fabric Properties for Virtual Simulation—A Critical Review”. In: IEEE SA INDUSTRY CONNECTIONS (2020), pp. 1–43 (cit. on p. 158).

[KL51] Solomon Kullback and Richard A Leibler. “On Information and Sufficiency”. In: The Annals of Mathematical Statistics 22.1 (1951), pp. 79–86 (cit. on pp. 196, 197).

[Kum+22] Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. “Fine-tuning Can Distort Pretrained Features and Underperform Out-of-distribution”. In: arXiv preprint arXiv:2202.10054 (2022) (cit. on p. 159).

[Kum+19] Manoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, and Durk Kingma. “Videoflow: A Flow-based Generative Model for Video”. In: arXiv preprint arXiv:1903.01434 2.5 (2019), p. 3 (cit. on p. 186).

[Kur+19] Karol Kurach, Mario Lučić, Xiaohua Zhai, Marcin Michalski, and Sylvain Gelly. “A Large-Scale Study on Regularization and Normalization in GANs”. In: International Conference on Machine Learning (ICML). PMLR. 2019, pp. 3581–3590 (cit. on pp. 77, 94).

[Kur+22] Alexander Kurz, Katja Hauser, Hendrik Alexander Mehrtens, Eva Krieghoff-Henning, Achim Hekler, Jakob Nikolas Kather, Stefan Fröhling, Christof von Kalle, Titus Josef Brinker, et al. “Uncertainty Estimation in Medical Image Classification: Systematic Review”. In: JMIR Medical Informatics 10.8 (2022) (cit. on p. 126).

[Kuz+21] Alexandr Kuznetsov, Krishna Mullia, Zexiang Xu, Miloš Hašan, and Ravi Ramamoorthi. “NeuMIP: Multi-Resolution Neural Materials”. In: ACM Transactions on Graphics (TOG) 40.4 (2021), pp. 1–13 (cit. on pp. 18, 102, 103, 106, 107, 109–116, 119, 120, 187, 207, 259, 264).

[Kuz+22] Alexandr Kuznetsov, Xuezheng Wang, Krishna Mullia, Fujun Luan, Zexiang Xu, Milos Hasan, and Ravi Ramamoorthi. “Rendering Neural Materials on Curved Surfaces”. In: ACM SIGGRAPH 2022 Conference Proceedings. 2022, pp. 1–9 (cit. on pp. 102, 104, 116, 187).

[Kwa+05] Vivek Kwatra, Irfan Essa, Aaron Bobick, and Nipun Kwatra. “Texture Optimization for Example-Based Synthesis”. In: ACM Transactions on Graphics (TOG). Vol. 24. 3. ACM. 2005, pp. 795–802 (cit. on pp. 72, 74).

[Kwa+03] Vivek Kwatra, Arno Schödl, Irfan Essa, Greg Turk, and Aaron Bobick. “Graphcut Textures: Image and Video Synthesis Ysing Graph Cuts”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 277–286 (cit. on p. 74).

[LGG19] Manuel Lagunas, Elena Garces, and Diego Gutierrez. “Learning Icons Appearance Similarity”. In: Multimedia Tools and Applications 78.8 (2019), pp. 10733–10751 (cit. on p. 160).

[Lag+19] Manuel Lagunas, Sandra Malpica, Ana Serrano, Elena Garces, Diego Gutierrez, and Belen Masia. “A Similarity Measure for Material Appearance”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–12 (cit. on pp. 157, 160, 180).

[LKA13] Samuli Laine, Tero Karras, and Timo Aila. “Megakernels Considered Harmful: Wavefront Path Tracing on GPUs”. In: Proceedings of the 5th High-Performance Graphics Conference. 2013, pp. 137–143 (cit. on pp. 109, 195).

[LPB17] Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. “Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on pp. 126, 130).

[LS98] Greg Ward Larson and Rob Shakespeare. Rendering with Radiance: The Art and Science of Lighting Visualization. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1998 (cit. on p. 185).

[Lav+21] Guillaume Lavoué, Nicolas Bonneel, Jean-Philippe Farrugia, and Cyril Soler. “Perceptual Quality of BRDF Approximations: Dataset and Metrics”. In: Computer Graphics Forum. Vol. 40. 2. Wiley Online Library. 2021, pp. 327–338 (cit. on pp. 130, 131, 193).

[Led+17] Christian Ledig, Lucas Theis, Ferenc Huszár, et al. “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2017-Janua (Sept. 2017), pp. 105–114 (cit. on pp. 79, 94).

[LH06] Sylvain Lefebvre and Hugues Hoppe. “Appearance-Space Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 25.3 (2006), pp. 541–548 (cit. on pp. 34, 104).

[Lei+21] Zhao Lei, Yi Zeng, Peng Liu, and Xiaohui Su. “Active Deep Learning for Hyperspectral Image Classification with Uncertainty Learning”. In: IEEE Geoscience and Remote Sensing Letters 19 (2021) (cit. on p. 126).

[LM01] Thomas Leung and Jitendra Malik. “Representing and Recognizing the Visual Appearance of Materials Using Three-Dimensional Textons”. In: International Journal of Computer Vision 43.1 (2001), pp. 29–44 (cit. on pp. 33, 39).

[LLB22] Wei-Hong Li, Xialei Liu, and Hakan Bilen. “Cross-domain Few-shot Learning with Task-specific Adapters”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, pp. 7161–7170 (cit. on p. 159).

[Li+17a] Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. “Modeling Surface Appearance from a Single Photograph Using Self-Augmented Convolutional Neural Networks”. In: ACM Transactions on Graphics (TOG) 36.4 (2017), pp. 1–11 (cit. on pp. 34, 124).

[Li+19] Xiaoyu Li, Bo Zhang, Pedro V Sander, and Jing Liao. “Blind Geometric Distortion Correction on Images through Deep Learning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4855–4864 (cit. on p. 93).

[Li+22] Yifei Li, Tao Du, Kui Wu, Jie Xu, and Wojciech Matusik. “DiffCloth: Differentiable Cloth Simulation with Dry Frictional Contact”. In: ACM Transactions on Graphics (TOG) (2022) (cit. on p. 158).

[Li+17b] Yijun Li, Chen Fang, Jimei Yang, Zhaowen Wang, Xin Lu, and Ming-Hsuan Yang. “Universal Style Transfer via Feature Transforms”. In: Advances in Neural Information Processing Systems. 2017, pp. 386–396 (cit. on p. 32).

[LAA08] Yuanzhen Li, Edward Adelson, and Aseem Agarwala. “ScribbleBoost: Adding Classification to Edge-Aware Interpolation of Local Image and Video Adjustments”. In: Computer Graphics Forum. Vol. 27. 4. Wiley Online Library. 2008, pp. 1255–1264 (cit. on p. 33).

[LS18] Zhengqi Li and Noah Snavely. “Learning Intrinsic Image Decomposition From Watching the World”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 78).

[Li+20] Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Man-mohan Chandraker. “Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF From a Single Image”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 2475–2484 (cit. on pp. 72, 74, 87–91, 95, 99, 114, 203).

[LSC18] Zhengqin Li, Kalyan Sunkavalli, and Manmohan Chandraker. “Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 72–87 (cit. on pp. 72, 125).

[Li+18] Zhengqin Li, Zexiang Xu, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. “Learning to Reconstruct Shape and Spatially-varying Reflectance from a Single Image”. In: ACM Transactions on Graphics (TOG) 37.6 (2018), pp. 1–11 (cit. on pp. 125, 129).

[LLK19] Junbang Liang, Ming Lin, and Vladlen Koltun. “Differentiable Cloth Simulation for Inverse Problems”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 158).

[Lia+17] Jing Liao, Yuan Yao, Lu Yuan, Gang Hua, and Sing Bing Kang. “Visual Attribute Transfer through Deep Image Analogy”. In: ACM Transactions on Graphics (TOG) 36.4 (2017), pp. 1–15 (cit. on pp. 31, 32, 46, 47, 56, 57, 60).

[Lic+21] Daniel Lichy, Jiaye Wu, Soumyadip Sengupta, and David W Jacobs. “Shape and Material Capture at Home”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 6123–6133 (cit. on p. 129).

[LPG19] Yiming Lin, Pieter Peers, and Abhijeet Ghosh. “On-Site Example-Based Material Appearance Acquisition”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 15–25 (cit. on pp. 34, 105).

[Liu+20] Guilin Liu, Rohan Taori, Ting-Chun Wang, Zhiding Yu, Shiqiu Liu, Fitsum A Reda, Karan Sapra, Andrew Tao, and Bryan Catanzaro. “Transposer: Universal Texture Synthesis Using Feature Maps as Transposed Convolution Filter”. In: arXiv preprint arXiv:2007.07243 (2020) (cit. on pp. 72, 74, 75, 92).

[Liu+19a] Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. “On the Variance of the Adaptive Learning Rate and Beyond”. In: arXiv preprint arXiv:1908.03265 (2019) (cit. on p. 110).

[Liu+19b] Ming-Yu Liu, Xun Huang, Arun Mallya, Tero Karras, Timo Aila, Jaakko Lehtinen, and Jan Kautz. “Few-shot Unsupervised Image-to-Image Translation”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019, pp. 10551–10560 (cit. on p. 37).

[Liu+21] Steven Liu, Xiuming Zhang, Zhoutong Zhang, Richard Zhang, Jun-Yan Zhu, and Bryan Russell. “Editing conditional radiance fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 5773–5783 (cit. on p. 116).

[Liu+22a] Yuan Liu, Sida Peng, Lingjie Liu, Qianqian Wang, Peng Wang, Christian Theobalt, Xiaowei Zhou, and Wenping Wang. “Neural Rays for Occlusion-aware Image-based Rendering”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 7824–7833 (cit. on p. 203).

[Liu+22b] Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. “A Convnet for the 2020s”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 11976–11986 (cit. on pp. 108, 109, 119).

[LSD15] Jonathan Long, Evan Shelhamer, and Trevor Darrell. “Fully Convolutional Networks for Semantic Segmentation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2015, pp. 3431–3440 (cit. on pp. 83, 94).

[LH17] Ilya Loshchilov and Frank Hutter. “Decoupled Weight Decay Regularization”. In: arXiv preprint arXiv:1711.05101 (2017) (cit. on p. 176).

[Lua+21] Fujun Luan, Shuang Zhao, Kavita Bala, and Zhao Dong. “Unified Shape and SVBRDF Recovery Using Differentiable Monte Carlo Rendering”. In: Computer Graphics Forum. Vol. 40. 4. Wiley Online Library. 2021, pp. 101–113 (cit. on pp. 125, 129).

[Lug+20] Andreas Lugmayr, Martin Danelljan, Luc Van Gool, and Radu Timofte. “Srflow: Learning the Super-resolution Space with Normalizing Flow”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 715–732 (cit. on p. 186).

[LM08] Christiane Luible and Nadia Magnenat-Thalmann. “The Simulation of Cloth Using Accurate Physical Parameters”. In: Proceedings of the Tenth IASTED International Conference on Computer Graphics and Imaging (CGIM ’08). 2008 (cit. on p. 158).

[LKS15] Zhaoliang Lun, Evangelos Kalogerakis, and Alla Sheffer. “Elements of Style: Learning Perceptual Shape Style Similarity”. In: ACM Transactions on graphics (TOG) 34.4 (2015), pp. 1–14 (cit. on p. 160).

[MHN13] Andrew L Maas, Awni Y Hannun, and Andrew Y Ng. “Rectifier Nonlinearities Improve Neural Network Acoustic Models”. In: International Conference on Machine Learning (ICML). Vol. 30. 1. Citeseer. 2013, p. 3 (cit. on pp. 77, 94, 176).

[Mag+07] Nadia Magnenat-Thalmann, Christiane Luible, Pascal Volino, and Etienne Lyard. “From Measured Fabric to the Simulation of Cloth”. In: 2007 10th IEEE International Conference on Computer-Aided Design and Computer Graphics. IEEE. 2007, pp. 7–18 (cit. on p. 158).

[MR10] Sébastien Marcel and Yann Rodriguez. “Torchvision the Machine-vision Package of Torch”. In: Proceedings of the 18th ACM International Conference on Multimedia. 2010, pp. 1485–1488 (cit. on pp. 67, 84, 110, 142, 194).

[Mar+20] Morteza Mardani, Guilin Liu, Aysegul Dundar, Shiqiu Liu, Andrew Tao, and Bryan Catanzaro. “Neural FFTs for Universal Texture Image Synthesis”. In: Advances in Neural Information Processing Systems 33 (2020) (cit. on pp. 72, 74, 75, 85, 107, 128).

[MFS08] Bruce A Maxwell, Richard M Friedhoff, and Casey A Smith. “A biilluminant dichromatic reflection model for understanding images”. In: 2008 IEEE Conference on Computer Vision and Pattern Recognition. IEEE. 2008, pp. 1–8 (cit. on p. 10).

[Maz+19] Ilya Mazlov, Sebastian Merzbach, Elena Trunz, and Reinhard Klein. “Neural Appearance Synthesis and Transfer”. In: (2019), pp. 35–39 (cit. on p. 31).

[MHM18] Leland McInnes, John Healy, and James Melville. “Umap: Uniform Manifold Approximation and Projection for Dimension Reduction”. In: arXiv preprint arXiv:1802.03426 (2018) (cit. on p. 133).

[Meh+21] Ishit Mehta, Michaël Gharbi, Connelly Barnes, Eli Shechtman, Ravi Ramamoorthi, and Manmohan Chandraker. “Modulated Periodic Activations for Generalizable Local Functional Representations”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 14214–14223 (cit. on pp. 109, 186, 203).

[MR21] Sachin Mehta and Mohammad Rastegari. “MobileViT: Light-Weight, General-Purpose, and Mobile-Friendly Vision Transformer”. In: International Conference on Learning Representations (ICLR). 2021 (cit. on pp. 124, 127, 132, 141, 144).

[Mel+12] Francho Melendez, Mashhuda Glencross, Jon Starck, and Gregory J Ward. “Transfer of Albedo and Local Depth Variation to Photo-Textures”. In: Proceedings of the 9th European Conference on Visual Media Production. 2012, pp. 40–48 (cit. on pp. 31, 34).

[Mel+21] Joe Mellor, Jack Turner, Amos Storkey, and Elliot J Crowley. “Neural Architecture Search Without Training”. In: International Conference on Machine Learning (ICML). PMLR. 2021, pp. 7588–7598 (cit. on p. 92).

[Mer+19] Sebastian Merzbach, Max Hermann, Martin Rump, and Reinhard Klein. “Learned Fitting of Spatially Varying BRDFs”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 193–205 (cit. on p. 30).

[MWK17] Sebastian Merzbach, Michael Weinmann, and Reinhard Klein. “High-quality Multi-Spectral Reflectance Acquisition with X-rite TAC7”. In: Proceedings of the Workshop on Material Appearance Modeling. 2017, pp. 11–16 (cit. on pp. 35, 39).

[MNG17] Lars Mescheder, Sebastian Nowozin, and Andreas Geiger. “Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks”. In: International Conference on Machine Learning (ICML). PMLR. 2017, pp. 2391–2400 (cit. on p. 126).

[Mic+18] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. “Mixed Precision Training”. In: International Conference on Learning Representations (ICLR). 2018 (cit. on pp. 56, 67, 84, 94, 110, 142, 176, 194).

[Mig+12] Eder Miguel, Derek Bradley, Bernhard Thomaszewski, Bernd Bickel, Wojciech Matusik, Miguel A Otaduy, and Steve Marschner. “Data-driven Estimation of Cloth Simulation Models”. In: Computer Graphics Forum. Vol. 31. 2pt2. Wiley Online Library. 2012, pp. 519–528 (cit. on p. 158).

[Mil+20] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2020 (cit. on pp. 186, 187, 207, 212, 262, 264).

[Min95] Pier Giorgio Minazio. “FAST–Fabric Assurance by Simple Testing”. In: International Journal of Clothing Science and Technology (1995) (cit. on pp. 156, 158).

[Miy+18] Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. “Spectral Normalization for Generative Adversarial Networks”. In: International Conference on Learning Representations. 2018 (cit. on pp. 141, 145).

[Mor+20] Alexander Mordvintsev, Ettore Randazzo, Eyvind Niklasson, and Michael Levin. “Growing Neural Cellular Automata”. In: Distill (2020). https://distill.pub/2020/growing-ca (cit. on pp. 15, 75, 257).

[Mor15] Andrew Morgan. The True Cost. 2015 (cit. on p. 2).

[Mor+17] Joep Moritz, Stuart James, Tom S.F. Haines, Tobias Ritschel, and Tim Weyrich. “Texture Stationarization: Turning Photos into Tileable Textures”. In: Eurographics Symposium on Geometry Processing. Vol. 36. 2. 2017, pp. 177–188 (cit. on pp. 15, 72–74, 88–90, 99, 114, 257).

[MMK03] Gero Müller, Jan Meseth, and Reinhard Klein. “Compression and Real-Time Rendering of Measured BTFs Using Local PCA.” In: VMV. 2003, pp. 271–279 (cit. on p. 103).

[Mül+22] Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. “Instant Neural Graphics Primitives with a Multiresolution Hash Encoding”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), 102:1–102:15 (cit. on pp. 203, 212).

[Mül+19] Thomas Müller, Brian McWilliams, Fabrice Rousselle, Markus Gross, and Jan Novák. “Neural Importance Sampling”. In: ACM Transactions on Graphics (TOG) 38.5 (2019), pp. 1–19 (cit. on pp. 116, 186, 191, 196, 197).

[Mül+20] Thomas Müller, Fabrice Rousselle, Alexander Keller, and Jan Novák. “Neural Control Variates”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), pp. 1–19 (cit. on pp. 186, 191).

[Mun+22] Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas Müller, and Sanja Fidler. “Extracting Triangular 3D Models, Materials, and Lighting From Images”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 8280–8290 (cit. on pp. 127, 129).

[Mur+20] J Krishna Murthy, Miles Macklin, Florian Golemo, Vikram Voleti, Linda Petrini, Martin Weiss, Breandan Considine, Jérôme Parent-Lévesque, Kevin Xie, Kenny Erleben, et al. “GradSim: Differentiable simulation for System Identification and Visuomotor Control”. In: International Conference on Learning Representations (ICLR). 2020 (cit. on p. 158).

[Naf+16] Hossein Ziaei Nafchi, Atena Shahkolaei, Rachid Hedjam, and Mohamed Cheriet. “Mean Deviation Similarity Index: Efficient and Reliable Full-reference Image Quality Evaluator”. In: IEEE Access 4 (2016), pp. 5579–5590 (cit. on p. 160).

[Nag+15] Koki Nagano, Graham Fyffe, Oleg Alexander, Jernej Barbic, Hao Li, Abhijeet Ghosh, and Paul E Debevec. “Skin Microstructure Deformation with Displacement Map Convolution.” In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 109–1 (cit. on p. 105).

[NH10] Vinod Nair and Geoffrey E Hinton. “Rectified Linear Units Improve Restricted Boltzmann Machines”. In: International Conference on Machine Learning (ICML). 2010 (cit. on pp. 54, 109, 111, 120).

[Nam+16] Giljoo Nam, Joo Ho Lee, Hongzhi Wu, Diego Gutierrez, and Min H Kim. “Simultaneous Acquisition of Microscale Reflectance and Normals”. In: ACM Transactions on Graphics (TOG) 35.6 (2016), pp. 1–11 (cit. on pp. 35, 39).

[NK18] Hyeonseob Nam and Hyo Eun Kim. “Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks”. In: Advances in Neural Information Processing Systems. Vol. 2018-Decem. 2018, pp. 2558–2567 (cit. on p. 77).

[NSO12] Rahul Narain, Armin Samii, and James F O’brien. “Adaptive Anisotropic Remeshing for Cloth Simulation”. In: ACM Transactions on Graphics (TOG) 31.6 (2012), pp. 1–10 (cit. on p. 162).

[ND06] Addy Ngan and Frédo Durand. “Statistical Acquisition of Texture Appearance”. In: Proceedings of the 17th Eurographics Conference on Rendering Techniques. 2006, pp. 31–40 (cit. on p. 63).

[NJR15] Jannik Boll Nielsen, Henrik Wann Jensen, and Ravi Ramamoorthi. “On Optimal, Minimal BRDF Sampling for Reflectance Acquisition”. In: ACM Transactions on Graphics (TOG) 34.6 (2015), pp. 1–11 (cit. on pp. 42, 116, 130).

[Nik+21] Eyvind Niklasson, Alexander Mordvintsev, Ettore Randazzo, and Michael Levin. “Self-Organising Textures”. In: Distill (2021). https://distill.pub/selforg/2021/textures (cit. on pp. 72, 74, 75, 89–91, 93, 95, 99).

[Nim+19] Merlin Nimier-David, Delio Vicini, Tizian Zeltner, and Wenzel Jakob. “Mitsuba 2: A Retargetable Forward and Inverse Renderer”. In: ACM Transactions on Graphics (TOG) 38.6 (2019), pp. 1–17 (cit. on pp. 6, 109, 202).

[ODO16] Augustus Odena, Vincent Dumoulin, and Chris Olah. “Deconvolution and Checker-board Artifacts”. In: Distill (2016) (cit. on p. 109).

[Ouy+21] Yaobin Ouyang, Shiqiu Liu, Markus Kettunen, Matt Pharr, and Jacopo Pantaleoni. “ReSTIR GI: Path Resampling for Real-time Path Tracing”. In: Computer Graphics Forum. Vol. 40. 8. 2021, pp. 17–29 (cit. on p. 184).

[Pap+21] George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. “Normalizing Flows for Probabilistic Modeling and Inference”. In: The Journal of Machine Learning Research 22.1 (2021), pp. 2617–2680 (cit. on p. 191).

[PPM17] George Papamakarios, Theo Pavlakou, and Iain Murray. “Masked Autoregressive Flow for Density Estimation”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 191).

[PSM19] George Papamakarios, David Sterratt, and Iain Murray. “Sequential Neural Likelihood: Fast Likelihood-free Inference with Autoregressive Flows”. In: The 22nd International Conference on Artificial Intelligence and Statistics. PMLR. 2019, pp. 837–848 (cit. on p. 191).

[Par+20] Taesung Park, Alexei A Efros, Richard Zhang, and Jun-Yan Zhu. “Contrastive Learning for Unpaired Image-to-Image Translation”. In: arXiv preprint arXiv:2007.15651 (2020) (cit. on pp. 37, 210).

[Pas+19] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. “Pytorch: An Imperative Style, High-performance Deep Learning Library”. In: Advances in Neural Information Processing Systems. Vol. 32. 2019 (cit. on pp. 39, 55, 67, 84, 94, 110, 142, 176, 194).

[Pha+20] Minh Quang Pham, Josep-Maria Crego, François Yvon, and Jean Senellart. “A study of Residual Adapters for Multi-domain Neural Machine Translation”. In: Conference on Machine Translation. 2020 (cit. on p. 159).

[Pha19] Matt Pharr. Visualizing Warping Strategies for Sampling Environment Map Lights. 2019 (cit. on pp. 189, 199, 200).

[PJH20] Matt Pharr, Wenzel Jakob, and Greg Humphreys. Implementation of the forthcoming 4th edition of Physically Based Rendering: From Theory to Implementation. 2020 (cit. on pp. 3, 109, 189, 200).

[PJH16] Matt Pharr, Wenzel Jakob, and Greg Humphreys. Physically Based Rendering: From Theory to Implementation. 3rd. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2016 (cit. on pp. 185, 189, 200, 201).

[PS00] Javier Portilla and Eero P Simoncelli. “A Parametric Texture Model Based on Joint Statistics of Complex Wavelet Coefficients”. In: International Journal of Computer Vision 40.1 (2000), pp. 49–70 (cit. on p. 74).

[Pow13] Jess Power. “Fabric Objective Measurements for Commercial 3D Virtual Garment Simulation”. In: International Journal of Clothing Science and Technology (2013) (cit. on p. 158).

[Pra+18] Ekta Prashnani, Hong Cai, Yasamin Mostofi, and Pradeep Sen. “Pieapp: Perceptual Image-error Assessment Through Pairwise Preference”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 1808–1817 (cit. on p. 160).

[Raa+18] Lara Raad, Axel Davy, Agnès Desolneux, and Jean-Michel Morel. “A Survey of Exemplar-Based Texture Synthesis”. In: Annals of Mathematical Sciences and Applications 3.1 (2018), pp. 89–148 (cit. on p. 73).

[Rad+21] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. “Learning Transferable Visual Models from Natural Language Supervision”. In: International Conference on Machine Learning (ICML). PMLR. 2021, pp. 8748–8763 (cit. on p. 159).

[Rag+19] Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio. “Transfusion: Under-standing Transfer Learning for Medical Imaging”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 159).

[Rai+22] Gilles Rainer, Adrien Bousseau, Tobias Ritschel, and George Drettakis. “Neural Pre-computed Radiance Transfer”. In: Computer Graphics Forum. Vol. 41. 2. Wiley Online Library. 2022, pp. 365–378 (cit. on pp. 105, 116, 203).

[Rai+20] Gilles Rainer, Abhijeet Ghosh, Wenzel Jakob, and Tim Weyrich. “Unified Neural Encod-ing of BTFs”. In: Computer Graphics Forum. Vol. 39. 2. The Eurographics Association. 2020, pp. 1–13 (cit. on pp. 18, 33, 102, 103, 109, 112, 113, 116, 259).

[Rai+19] Gilles Rainer, Wenzel Jakob, Abhijeet Ghosh, and Tim Weyrich. “Neural BTF Compres-sion and Interpolation”. In: Computer Graphics Forum. Vol. 38. 2. Wiley Online Library. 2019, pp. 235–244 (cit. on pp. 18, 33, 39, 40, 102, 103, 106, 109, 112, 113, 259).

[Ram+19] Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. “Stand-alone Self-attention in Vision Models”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 164).

[RH01a] Ravi Ramamoorthi and Pat Hanrahan. “An Efficient Representation for Irradiance Environment Maps”. In: Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques. SIGGRAPH ’01. New York, NY, USA: Association for Computing Machinery, 2001, pp. 497–500 (cit. on p. 13).

[RH01b] Ravi Ramamoorthi and Pat Hanrahan. “An Efficient Representation for Irradiance Environment Maps”. In: Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques. 2001, pp. 497–500 (cit. on pp. 196, 199, 202).

[Ras+20] Abdullah Haroon Rasheed, Victor Romero, Florence Bertails-Descoubes, Stefanie Wuhrer, Jean-Sébastien Franco, and Arnaud Lazarus. “Learning to Measure the Static Friction Coefficient in Cloth Contact”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 9912–9921 (cit. on p. 159).

[RD17] Ajita Rattani and Reza Derakhshani. “On Fine-tuning Convolutional Neural Networks for Smartphone Based Ocular Recognition”. In: 2017 IEEE international joint conference on biometrics (IJCB). IEEE. 2017, pp. 762–767 (cit. on p. 159).

[RBV18] Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. “Efficient Parametrization of Multi-domain Deep Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8119–8127 (cit. on p. 159).

[RBV17] Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. “Learning Multiple Visual Domains with Residual Adapters”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 159).

[RJ19] A Sai Bharadwaj Reddy and D Sujitha Juliet. “Transfer Learning with ResNet-50 for Malaria Cell-image Classification”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). IEEE. 2019, pp. 0945–0949 (cit. on p. 159).

[Ree+15] Scott E Reed, Yi Zhang, Yuting Zhang, and Honglak Lee. “Deep Visual Analogy-Making”. In: Advances in Neural Information Processing Systems. 2015, pp. 1252–1260 (cit. on p. 32).

[Rei+14] Klein Reinhard, Rump Martin, Weinmann Michael, Sarlette Ralf, and Schwartz Christopher. “Design and Implementation of Practical Bidirectional Texture Function Measurement Devices focusing on the Developments at the University of Bonn”. In: Sensors 14.5 (2014) (cit. on pp. 109, 114, 117).

[Rei+18] Rafael Reisenhofer, Sebastian Bosse, Gitta Kutyniok, and Thomas Wiegand. “A Haar Wavelet-based Perceptual Similarity Index for Image Quality Assessment”. In: Signal Processing: Image Communication 61 (2018), pp. 33–43 (cit. on p. 160).

[RSS16] Nathalie Remy, Eveline Speelman, and Steven Swartz. Style that’s Sustainable: A New Fast-Fashion Formula. 2016 (cit. on p. 2).

[Ren+17] Jian Ren, Xiaohui Shen, Zhe Lin, Radomir Mech, and David J Foran. “Personalized Image Aesthetics”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 638–647 (cit. on p. 74).

[Rez+20] Danilo Jimenez Rezende, George Papamakarios, Sébastien Racaniere, Michael Albergo, Gurtej Kanwar, Phiala Shanahan, and Kyle Cranmer. “Normalizing Flows on Tori and Spheres”. In: International Conference on Machine Learning (ICML). PMLR. 2020, pp. 8083–8092 (cit. on p. 189).

[Al-+16] Rami Al-Rfou, Guillaume Alain, Amjad Almahairi, Christof Angermueller, Dzmitry Bahdanau, Nicolas Ballas, Frédéric Bastien, Justin Bayer, Anatoly Belikov, Alexander Belopolsky, et al. “Theano: A Python Framework for Fast Computation of Mathematical Expressions”. In: arXiv e-prints (2016), arXiv–1605 (cit. on p. 95).

[Rho+22] Daniel Rho, Junwoo Cho, Jong Hwan Ko, and Eunbyung Park. “Neural Residual Flow Fields for Efficient Video Representations”. In: Proceedings of the Asian Conference on Computer Vision. 2022, pp. 3447–3463 (cit. on p. 187).

[Rib+20] Edgar Riba, Dmytro Mishkin, Daniel Ponsa, Ethan Rublee, and Gary Bradski. “Kornia: An Open Source Differentiable Computer Vision Library for Pytorch”. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2020, pp. 3674–3683 (cit. on pp. 67, 111, 142, 176).

[Ric+16] Stephan R. Richter, Vibhav Vineet, Stefan Roth, and Vladlen Koltun. “Playing for Data: Ground Truth from Computer Games”. In: European Conference on Computer Vision (ECCV). Ed. by Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling. Vol. 9906. LNCS. Springer International Publishing, 2016, pp. 102–118 (cit. on p. 212).

[RPG16] Jérémy Riviere, Pieter Peers, and Abhijeet Ghosh. “Mobile Surface Reflectometry”. In: Computer Graphics Forum. Vol. 35. 1. Wiley Online Library. 2016, pp. 191–202 (cit. on pp. 31, 34).

[Rod+23a] Carlos Rodriguez-Pardo, Henar Dominguez-Elvira, David Pascual-Hernandez, and Elena Garces. “UMat: Uncertainty-Aware Single Image High Resolution Material Capture”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023) (cit. on pp. 24, 62, 76, 102, 107, 108, 110, 116, 193, 268).

[RG21] Carlos Rodriguez-Pardo and Elena Garces. “Neural Photometry-guided Visual Attribute Transfer”. In: IEEE Transactions on Visualization and Computer Graphics (2021) (cit. on pp. 23, 25, 63–70, 102, 105, 107, 108, 110, 114, 125, 129, 131, 142, 193, 268).

[RG22] Carlos Rodriguez-Pardo and Elena Garces. “SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps”. In: IEEE Transactions on Visualization and Computer Graphics (2022) (cit. on pp. 23, 25, 64, 67, 69, 104, 107, 108, 110, 114, 127, 128, 141, 268).

[Rod+23b] Carlos Rodriguez-Pardo, Melania Prieto-Martin, Dan Casas, and Elena Garces. “How Will It Drape Like? Capturing Fabric Mechanics from Depth Images”. In: Computer Graphics Forum (Proc. of Eurographics) (2023) (cit. on pp. 18, 23, 212, 268).

[Rod+19] Carlos Rodriguez-Pardo, Sergio Suja, David Pascual, Jorge Lopez-Moreno, and Elena Garces. “Automatic Extraction and Synthesis of Regular Repeatable Patterns”. In: Computers & Graphics 83 (2019), pp. 33–41 (cit. on pp. 15, 34, 72, 75, 79, 83, 87, 89–91, 93, 95, 99, 114, 257).

[RB19] Carlos Rodrıguez-Pardo and Hakan Bilen. “Personalised Aesthetics with Residual Adapters”. In: Iberian Conference on Pattern Recognition and Image Analysis. Springer. 2019, pp. 508–520 (cit. on pp. 74, 90, 159).

[Rom+22] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. “High-Resolution Image Synthesis with Latent Diffusion Models”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 10684–10695 (cit. on pp. 127, 209, 210).

[RFB15] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. “U-Net: Convolutional Networks for Biomedical Image Segmentation”. In: International Conference on Medical image computing and computer-assisted intervention. Springer. 2015, pp. 234–241 (cit. on pp. 37–39, 54, 55, 64, 66, 67, 69, 78, 108, 111, 124, 127, 132, 141).

[RSK13] Roland Ruiters, Christopher Schwartz, and Reinhard Klein. “Example-based Interpolation and Synthesis of Bidirectional Texture Functions”. In: Computer Graphics Forum. Vol. 32. 2pt3. Wiley Online Library. 2013, pp. 361–370 (cit. on p. 63).

[RSK10] Martin Rump, Ralf Sarlette, and Reinhard Klein. “Groundtruth Data for Multispectral Bidirectional Texture Functions”. In: Conference on Colour in Graphics, Imaging, and Vision. Vol. 2010. 1. Society for Imaging Science and Technology. 2010, pp. 326–331 (cit. on p. 116).

[Run+20] Tom FH Runia, Kirill Gavrilyuk, Cees GM Snoek, and Arnold WM Smeulders. “Cloth in the Wind: A Case Study of Physical Measurement Through Simulation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 10498–10507 (cit. on p. 158).

[Rus+04] Daniel B Russakoff, Carlo Tomasi, Torsten Rohlfing, and Calvin R Maurer. “Image Similarity Using Mutual Information of Regions”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2004, pp. 596–607 (cit. on p. 132).

[Sah+22] Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. “Palette: Image-to-Image Diffusion Models”. In: ACM SIGGRAPH 2022 Conference Proceedings. 2022, pp. 1–10 (cit. on pp. 127, 209).

[San+19] Veit Sandfort, Ke Yan, Perry J Pickhardt, and Ronald M Summers. “Data Augmentation Using Generative Adversarial Networks (CycleGAN) to Improve Generalizability in CT Segmentation Tasks”. In: Scientific reports 9.1 (2019), pp. 1–9 (cit. on p. 45).

[SC20] Shen Sang and Manmohan Chandraker. “Single-Shot Neural Relighting and SVBRDF Estimation”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 85–101 (cit. on p. 125).

[SOC22] Igor Santesteban, Miguel A Otaduy, and Dan Casas. “SNUG: Self-Supervised Neural Dynamic Garments”. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022) (cit. on pp. 17, 259).

[SSK03] Mirko Sattler, Ralf Sarlette, and Reinhard Klein. “Efficient and Realistic Visualization of Cloth”. In: Rendering techniques. 2003, pp. 167–178 (cit. on p. 109).

[SSK20] Edgar Schonfeld, Bernt Schiele, and Anna Khoreva. “A U-Net Based Discriminator for Generative Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 8207–8216 (cit. on pp. 83, 93, 124, 128, 132, 141, 142).

[Sch+17] Vincent Schüssler, Eric Heitz, Johannes Hanika, and Carsten Dachsbacher. “Microfacet-based Normal Mapping for Robust Monte Carlo Path Tracing”. In: ACM Transactions on Graphics (TOG) 36.6 (2017), pp. 1–12 (cit. on p. 201).

[See66] Robert T Seeley. “Spherical Harmonics”. In: The American Mathematical Monthly 73.4P2 (1966), pp. 115–121 (cit. on pp. 185, 196).

[Sel+20] Raghavendra Selvan, Frederik Faye, Jon Middleton, and Akshay Pai. “Uncertainty Quantification in Medical Image Segmentation with Normalizing Flows”. In: Machine Learning in Medical Imaging. 2020, pp. 80–90 (cit. on p. 186).

[SKK18] Murat Sensoy, Lance Kaplan, and Melih Kandemir. “Evidential Deep Learning to Quantify Classification Uncertainty”. In: Advances in Neural Information Processing Systems 31 (2018) (cit. on pp. 126, 130).

[SDM19] Tamar Rott Shaham, Tali Dekel, and Tomer Michaeli. “SinGAN: Learning a Generative Model From a Single Natural Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on pp. 32, 38, 57, 75, 85, 86, 90).

[Shi+20] Liang Shi, Beichen Li, Miloš Hašan, Kalyan Sunkavalli, Tamy Boubekeur, Radomir Mech, and Wojciech Matusik. “Match: Differentiable Material Graphs for Procedural Material Capture”. In: ACM Transactions on Graphics (TOG) (2020) (cit. on pp. 76, 125, 137, 138, 143, 151–153).

[Sho+19] Assaf Shocher, Shai Bagon, Phillip Isola, and Michal Irani. “InGAN: Capturing and Retargeting the "DNA" of a Natural Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on p. 75).

[SCI18] Assaf Shocher, Nadav Cohen, and Michal Irani. ““Zero-Shot” Super-Resolution Using Deep Internal Learning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 75).

[SK19] Connor Shorten and Taghi M Khoshgoftaar. “A Survey on Image Data Augmentation for Deep Learning”. In: Journal of Big Data 6.1 (2019), p. 60 (cit. on pp. 36, 45).

[SZ15] Karen Simonyan and Andrew Zisserman. “Very Deep Convolutional Networks for Large-Scale Image Recognition”. In: International Conference on Learning Representations (ICLR). 2015 (cit. on pp. 32, 164).

[Sin+21] Abhishek Sinha, Kumar Ayush, Jiaming Song, Burak Uzkent, Hongxia Jin, and Stefano Ermon. “Negative Data Augmentation”. In: arXiv preprint arXiv:2102.05113 (2021) (cit. on p. 93).

[Sit+20] Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. “Implicit Neural Representations with Periodic Activation Functions”. In: Advances in Neural Information Processing Systems 33 (2020) (cit. on pp. 75, 102, 109, 112, 120, 186, 193, 194, 262).

[Sit+21] Vincent Sitzmann, Semon Rezchikov, William T. Freeman, Joshua B. Tenenbaum, and Fredo Durand. “Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering”. In: Advances in Neural Information Processing Systems. 2021 (cit. on p. 187).

[SKS02] Peter-Pike Sloan, Jan Kautz, and John Snyder. “Precomputed Radiance Transfer for Real-Time Rendering in Dynamic, Low-Frequency Lighting Environments”. In: ACM Transactions on Graphics (TOG) 21.3 (July 2002), pp. 527–536 (cit. on p. 13).

[Sne17] Xavier Snelgrove. “High-resolution Multi-scale Neural Texture Synthesis”. In: SIG-GRAPH Asia 2017 Technical Briefs. 2017, pp. 1–4 (cit. on p. 74).

[Sol+21] Ava P Soleimany, Alexander Amini, Samuel Goldman, Daniela Rus, Sangeeta N Bhatia, and Connor W Coley. “Evidential Deep Learning for Guided Molecular Property Prediction and Discovery”. In: ACS Central Science 7.8 (2021), pp. 1356–1367 (cit. on pp. 126, 135).

[Spe+22] Georg Sperl, Rosa M Sánchez-Banderas, Manwen Li, Chris Wojtan, and Miguel A Otaduy. “Estimation of Yarn-level Simulation Models for Production Fabrics”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), pp. 1–15 (cit. on p. 163).

[Sri+20] Pratul P Srinivasan, Ben Mildenhall, Matthew Tancik, Jonathan T Barron, Richard Tucker, and Noah Snavely. “Lighthouse: Predicting Lighting Volumes for Spatially-Coherent Illumination”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 8080–8089 (cit. on p. 203).

[Sri+14] Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”. In: The Journal of Machine Learning Research 15.1 (2014), pp. 1929–1958 (cit. on pp. 110, 127, 141, 176).

[Ste+14] Heinz Christian Steinhausen, Dennis den Brok, Matthias B Hullin, and Reinhard Klein. “Acquiring Bidirectional Texture Functions for Large-Scale Material Samples”. In: (2014) (cit. on pp. 30, 63).

[Ste+15a] Heinz Christian Steinhausen, Dennis den Brok, Matthias B. Hullin, and Reinhard Klein. “Extrapolating Large-Scale Material BTFs under Cross-Device Constraints”. In: Vision, Modeling & Visualization. Ed. by David Bommes, Tobias Ritschel, and Thomas Schultz. The Eurographics Association, 2015, pp. 143–150 (cit. on pp. 33, 104).

[Ste+15b] Heinz Christian Steinhausen, Rodrigo Martín, Dennis den Brok, Matthias B. Hullin, and Reinhard Klein. “Extrapolation of Bidirectional Texture Functions using Texture Synthesis guided by Photometric Normals”. In: Measuring, Modeling, and Reproducing Material Appearance II (SPIE 9398). Vol. 9398. 14. San Francisco, USA, Feb. 2015 (cit. on pp. 33, 104).

[Str+22] Yannick Strümpler, Janis Postels, Ren Yang, Luc Van Gool, and Federico Tombari. “Implicit Neural Representations for Image Compression”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 74–91 (cit. on p. 187).

[SGK21] Vadim Sushko, Juergen Gall, and Anna Khoreva. “One-Shot GAN: Learning to Generate Samples from Single Images and Videos”. In: arXiv preprint arXiv:2103.13389 (2021) (cit. on p. 75).

[SB08] Cédric Syllebranque and Samuel Boivin. “Estimation of Mechanical Parameters of Deformable Solids from Videos”. In: The Visual Computer 24.11 (2008), pp. 963–972 (cit. on p. 158).

[Szt+21] Alejandro Sztrajman, Gilles Rainer, Tobias Ritschel, and Tim Weyrich. “Neural BRDF Representation and Importance Sampling”. In: Computer Graphics Forum. Vol. 40. 6. Wiley Online Library. 2021, pp. 332–346 (cit. on pp. 105, 111, 116, 186, 187, 193, 203, 210).

[Tak+21] Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. “Neural Geometric Level of Detail: Real-time Rendering with Implicit 3D Shapes”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021 (cit. on p. 186).

[Tam+11] Omer Tamuz, Ce Liu, Serge Belongie, Ohad Shamir, and Adam Tauman Kalai. “Adaptively Learning the Crowd Kernel”. In: International Conference on Machine Learning (ICML). 2011, pp. 673–680 (cit. on pp. 170, 171).

[Tan+22] Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul Srinivasan, Jonathan T. Barron, and Henrik Kretzschmar. “Block-NeRF: Scalable Large Scene Neural View Synthesis”. In: arXiv. 2022 (cit. on p. 187).

[Tan+20] Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. “Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains”. In: Advances in Neural Information Processing Systems (2020) (cit. on pp. 75, 186, 193).

[Tew+22] Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, Wang Yifan, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, et al. “Advances in Neural Rendering”. In: Computer Graphics Forum. Vol. 41. 2. Wiley Online Library. 2022, pp. 703–735 (cit. on pp. 187, 212).

[Tex+20a] Ondřej Texler, David Futschik, Jakub Fišer, Michal Lukáč, Jingwan Lu, Eli Shechtman, and Daniel Sy`kora. “Arbitrary Style Transfer using Neurally-Guided Patch-Based Ssynthesis”. In: Computers & Graphics 87 (2020), pp. 62–71 (cit. on pp. 39, 45).

[Tex+20b] Ondřej Texler, David Futschik, Michal Kučera, Ondřej Jamriška, Šárka Sochorová, Menclei Chai, Sergey Tulyakov, and Daniel Sy`kora. “Interactive Video Stylization Using Few-Shot Patch-Based Training”. In: ACM Transactions on Graphics (TOG) 39.4 (2020), pp. 73–1 (cit. on pp. 30, 31, 33, 37, 39, 46, 53, 58, 61, 107, 108, 128, 129).

[TOB15] Lucas Theis, Aäron van den Oord, and Matthias Bethge. “A Note on the Evaluation of Generative Models”. In: arXiv preprint arXiv:1511.01844 (2015) (cit. on p. 41).

[Tom94] Shoji Tominaga. “Dichromatic reflection models for a variety of materials”. In: Color Research & Application 19.4 (1994), pp. 277–285 (cit. on p. 10).

[Ton+02] Xin Tong, Jingdan Zhang, Ligang Liu, Xi Wang, Baining Guo, and Heung-Yeung Shum. “Synthesis of Bidirectional Texture Functions on Arbitrary Surfaces”. In: ACM Transactions on Graphics (TOG) 21.3 (2002), pp. 665–672 (cit. on p. 103).

[TS67] Kenneth E Torrance and Ephraim M Sparrow. “Theory for Off-Specular Reflection from Roughened Surfaces”. In: Josa 57.9 (1967), pp. 1105–1114 (cit. on p. 11).

[Tu+20] Peihan Tu, Li-Yi Wei, Koji Yatani, Takeo Igarashi, and Matthias Zwicker. “Continuous Curve Textures”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), pp. 1–16 (cit. on p. 72).

[TK74] Amos Tversky and Daniel Kahneman. “Judgment under Uncertainty: Heuristics and Biases: Biases in Judgments Reveal Some Heuristics of Thinking Under Uncertainty.” In: Science 185.4157 (1974), pp. 1124–1131 (cit. on p. 169).

[Uly+16] Dmitry Ulyanov, Vadim Lebedev, Andrea Vedaldi, and Victor S Lempitsky. “Texture Networks: Feed-Forward Synthesis of Textures and Stylized Images.” In: International Conference of Machine Learning (ICML). Vol. 1. 2. 2016, p. 4 (cit. on p. 76).

[UVL18] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. “Deep Image Prior”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 75).

[UVL16] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. “Instance Normalization: The Missing Ingredient for Fast Stylization”. In: arXiv preprint arXiv:1607.08022 (2016) (cit. on pp. 77, 94).

[Van11] Dietger G Van Antwerpen. “Unbiased Physically Nased Rendering on the GPU”. MA thesis. Electrical Engineering, Mathematics and Computer Science, 2011 (cit. on p. 195).

[Vas+17] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. “Attention Is All You Need”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 164).

[VG95] Eric Veach and Leonidas J. Guibas. “Optimally Combining Sampling Techniques for Monte Carlo Rendering”. In: Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques. SIGGRAPH ’95. Association for Computing Machinery, 1995 (cit. on p. 184).

[VPS21] Giuseppe Vecchio, Simone Palazzo, and Concetto Spampinato. “SurfaceNet: Adversarial SVBRDF Estimation from a Single Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 12840–12848 (cit. on pp. 17, 76, 123, 125, 128, 129, 132, 258).

[Ver+22] Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T. Barron, and Pratul P. Srinivasan. “Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2022 (cit. on pp. 187, 212).

[VZH20] Yael Vinker, Nir Zabari, and Yedid Hoshen. “Training End-to-end Single Image Generators without GANs”. In: arXiv preprint arXiv:2004.06014 (2020) (cit. on p. 75).

[VMF09] Pascal Volino, Nadia Magnenat-Thalmann, and Francois Faure. “A Simple Approach to Nonlinear Tensile Stiffness for Accurate Cloth Simulation”. In: ACM Transactions on Graphics (TOG) 28.4 (2009), Article–No (cit. on pp. 158, 161, 162).

[Wal+14] Ingo Wald, Sven Woop, Carsten Benthin, Gregory S. Johnson, and Manfred Ernst. “Embree: A Kernel Framework for Efficient CPU Ray Tracing”. In: ACM Transactions on Graphics (TOG) 33.4 (2014) (cit. on p. 195).

[Wal+07] Bruce Walter, Stephen R Marschner, Hongsong Li, and Kenneth E Torrance. “Microfacet Models for Refraction through Rough Surfaces”. In: Proceedings of the 18th Eurographics Conference on Rendering Techniques. 2007, pp. 195–206 (cit. on p. 127).

[WZY22] Cairong Wang, Yiming Zhu, and Chun Yuan. “Diverse Image Inpainting with Normalizing Flow”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 53–69 (cit. on p. 186).

[Wan+22a] Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. “Clip-nerf: Text-and-image Driven Manipulation of Neural Radiance Fields”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 3835–3844 (cit. on p. 116).

[Wan+22b] Chen Wang, Xiang Wang, Jiawei Zhang, Liang Zhang, Xiao Bai, Xin Ning, Jun Zhou, and Edwin Hancock. “Uncertainty Estimation for Stereo Matching Based on Evidential Deep Learning”. In: Pattern Recognition 124 (2022), p. 108498 (cit. on p. 130).

[Wan+18] Guotai Wang, Wenqi Li, Maria A Zuluaga, Rosalind Pratt, Premal A Patel, Michael Aertsen, Tom Doel, Anna L David, Jan Deprest, Sébastien Ourselin, et al. “Interactive Medical Image Segmentation using Deep Learning with Image-specific Fine Tuning”. In: IEEE transactions on medical imaging 37.7 (2018), pp. 1562–1573 (cit. on p. 159).

[WOR11] Huamin Wang, James F O’Brien, and Ravi Ramamoorthi. “Data-driven Elastic Models for Cloth: Modeling and Measurement”. In: ACM Transactions on Graphics (TOG) 30.4 (2011), pp. 1–12 (cit. on p. 158).

[Wan+09] Jiaping Wang, Peiran Ren, Minmin Gong, John Snyder, and Baining Guo. “All-Frequency Rendering of Dynamic, Spatially-Varying Reflectance”. In: 28.5 (Dec. 2009), pp. 1–10 (cit. on p. 13).

[Wan+22c] Jiayi Wang, Diogo Luvizon, Franziska Mueller, Florian Bernard, Adam Kortylewski, Dan Casas, and Christian Theobalt. “HandFlow: Quantifying View-Dependent 3D Ambiguity in Two-Hand Reconstruction with Normalizing Flow”. In: International Symposium on Vision, Modeling, and Visualization. 2022 (cit. on p. 186).

[Wan+20a] Jiayun Wang, Yubei Chen, Rudrasis Chakraborty, and Stella X. Yu. “Orthogonal Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2020 (cit. on p. 93).

[Wan+22d] Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. “Fourier Plenoctrees for Dynamic Radiance Field Rendering in Real-time”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 13524–13534 (cit. on pp. 203, 212).

[Wan+20b] Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. “Linformer: Self-attention with Linear Complexity”. In: arXiv preprint arXiv:2006.04768 (2020) (cit. on pp. 124, 127, 141, 144).

[Wan+19] Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Bryan Catanzaro, and Jan Kautz. “Few-shot Video-to-Video Synthesis”. In: Advances in Neural Information Processing Systems. 2019, pp. 5013–5024 (cit. on p. 37).

[Wan+20c] Xiaofei Wang, Yiwen Han, Victor CM Leung, Dusit Niyato, Xueqiang Yan, and Xu Chen. “Convergence of Edge Computing and Deep Learning: A Comprehensive Survey”. In: IEEE Communications Surveys & Tutorials 22.2 (2020), pp. 869–904 (cit. on p. 212).

[Wan+20d] Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni. “Generalizing From a Few Examples: A Survey on Few-Shot Learning”. In: ACM Computing Surveys (CSUR) 53.3 (2020), pp. 1–34 (cit. on p. 37).

[Wan+04] Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. “Image Quality Assessment: From Error Visibility to Structural Similarity”. In: IEEE Transactions on Image Processing 13.4 (2004), pp. 600–612 (cit. on pp. 50, 54, 85, 86, 90, 92, 160, 198, 199, 202).

[Wan+21] Zian Wang, Jonah Philion, Sanja Fidler, and Jan Kautz. “Learning Indoor Inverse Rendering with 3D Spatially-Varying Lighting”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 12538–12547 (cit. on p. 203).

[Wan+23] Zian Wang, Tianchang Shen, Jun Gao, Shengyu Huang, Jacob Munkberg, Jon Hasselgren, Zan Gojcic, Wenzheng Chen, and Sanja Fidler. “Neural Fields meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2023 (cit. on pp. 203, 212).

[WGK14] Michael Weinmann, Juergen Gall, and Reinhard Klein. “Material Classification Based on Training Data Synthesized Using a BTF Database”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer International Publishing, 2014, pp. 156–171 (cit. on pp. 50, 103, 108, 111, 151).

[WKW16] Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang. “A Survey of Transfer Learning”. In: Journal of Big data 3.1 (2016), pp. 1–40 (cit. on p. 159).

[Wen+22] Tao Wen, Beibei Wang, Lei Zhang, Jie Guo, and Nicolas Holzschuch. “SVBRDF Recovery from a Single Image with Highlights Using a Pre-trained Generative Adversarial Network”. In: Computer Graphics Forum. Wiley Online Library. 2022 (cit. on p. 125).

[Woo+18] Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. “Cbam: Convolutional block attention module”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 3–19 (cit. on pp. 109, 119, 128, 141, 145, 164).

[WTX22] Qiling Wu, Jianchao Tan, and Kun Xu. “PaletteNeRF: Palette-based Color Editing for NeRFs”. In: arXiv preprint arXiv:2212.12871 (2022) (cit. on p. 116).

[WH18] Yuxin Wu and Kaiming He. “Group Normalization”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 3–19 (cit. on pp. 64, 67, 69, 127, 141, 144).

[Xie+22] Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar. “Neural Fields in Visual Computing and Beyond”. In: Computer Graphics Forum (2022) (cit. on pp. 187, 212).

[Xu+13] Kun Xu, Wei-Lun Sun, Zhao Dong, Dan-Yong Zhao, Run-Dong Wu, and Shi-Min Hu. “Anisotropic Spherical Gaussians”. In: ACM Transactions on Graphics (TOG) 32.6 (2013), pp. 1–11 (cit. on pp. 185, 196, 199, 202).

[Yan+21] Guandao Yang, Serge Belongie, Bharath Hariharan, and Vladlen Koltun. “Geometry processing with Neural Fields”. In: Advances in Neural Information Processing Systems. Vol. 34. 2021, pp. 22483–22497 (cit. on p. 186).

[YLL17] Shan Yang, Junbang Liang, and Ming C Lin. “Learning-based Cloth Material Recovery from Video”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2017, pp. 4383–4393 (cit. on pp. 156, 159).

[YL15] Shan Yang and Ming C Lin. “Materialcloning: Acquiring Elasticity Parameters from Images for Medical Applications”. In: IEEE Transactions on Visualization and Computer Graphics 22.9 (2015), pp. 2122–2135 (cit. on p. 158).

[Yan+18] Shan Yang, Zherong Pan, Tanya Amert, Ke Wang, Licheng Yu, Tamara Berg, and Ming C Lin. “Physics-inspired Garment Recovery from a Single-view Image”. In: ACM Transactions on Graphics (TOG) 37.5 (2018), pp. 1–14 (cit. on p. 158).

[Yao+22] Yao Yao, Jingyang Zhang, Jingbo Liu, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, and Long Quan. “NeILF: Neural Incident Light Field for Material and Lighting Estimation”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2022 (cit. on p. 187).

[Ye+22] Weicai Ye, Shuo Chen, Chong Bao, Hujun Bao, Marc Pollefeys, Zhaopeng Cui, and Guofeng Zhang. “Intrinsicnerf: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis”. In: arXiv preprint arXiv:2210.00647 (2022) (cit. on p. 116).

[Ye+21] Wenjie Ye, Yue Dong, Pieter Peers, and Baining Guo. “Deep Reflectance Scanning: Recovering Spatially-varying Material Appearance from a Flash-lit Video Sequence”. In: Computer Graphics Forum. Vol. 40. 6. Wiley Online Library. 2021, pp. 409–427 (cit. on p. 125).

[Ye+18] Wenjie Ye, Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. “Single Image Surface Appearance Modeling with Self-Augmented CNNs and Inexact Supervision”. In: Computer Graphics Forum. Vol. 37. 7. Wiley Online Library. 2018, pp. 201–211 (cit. on pp. 34, 125).

[Yep20] Tyler Yep. torchinfo. Mar. 2020 (cit. on p. 113).

[YS19] Ye Yu and William AP Smith. “InverseRenderNet: Learning Single Image Inverse Rendering”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 3155–3164 (cit. on pp. 78, 203).

[YS21] Ye Yu and William AP Smith. “Outdoor Inverse Rendering from a Single Image using Multiview Self-supervision”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 44.7 (2021), pp. 3659–3675 (cit. on p. 203).

[Yun+19] Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. “Cutmix: Regularization strategy to train strong classifiers with localizable features”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 6023–6032 (cit. on p. 93).

[Zha+19a] Bo Zhang, Mingming He, Jing Liao, Pedro V Sander, Lu Yuan, Amine Bermak, and Dong Chen. “Deep Exemplar-based Video Colorization”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 8052–8061 (cit. on p. 33).

[Zha+19b] Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. “Self-attention Generative Adversarial Networks”. In: International Conference on Machine Learning (ICML). PMLR. 2019, pp. 7354–7363 (cit. on pp. 164, 166, 176).

[ZDN16] Hang Zhang, Kristin Dana, and Ko Nishino. “Friction from Reflectance: Deep Reflectance Codes for Predicting Physical Surface Properties from One-shot In-field Reflectance”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2016, pp. 808–824 (cit. on p. 159).

[Zha+20] Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. “Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity”. In: International Conference on Learning Representations (ICLR). 2020 (cit. on pp. 110, 194).

[Zha+21] Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. “Physg: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 5453–5462 (cit. on p. 203).

[ZL12] Lin Zhang and Hongyu Li. “SR-SIM: A Fast and High Performance IQA Index Based on Spectral Residual”. In: IEEE International Conference on Image Processing. IEEE. 2012, pp. 1473–1476 (cit. on p. 160).

[ZSL14] Lin Zhang, Ying Shen, and Hongyu Li. “VSI: A Visual Saliency-induced Index for Perceptual Image Quality Assessment”. In: IEEE Transactions on Image processing 23.10 (2014), pp. 4270–4281 (cit. on p. 160).

[Zha+11] Lin Zhang, Lei Zhang, Xuanqin Mou, and David Zhang. “FSIM: A Feature Similarity Index for Image Quality Assessment”. In: IEEE Transactions on Image Processing 20.8 (2011), pp. 2378–2386 (cit. on p. 160).

[Zha+18] Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 586–595 (cit. on pp. 32, 39, 41, 49, 50, 54, 57, 67, 85, 86, 90, 110, 128, 142, 157, 160, 169, 170, 178, 198, 199, 202).

[ZJK20] Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. “Exploring Self-attention for Image Recognition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 10076–10085 (cit. on p. 164).

[Zho+20] Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. “Random Erasing Data Augmentation”. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 2020 (cit. on pp. 111, 129, 142, 176).

[Zho+16] Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. “Learning Deep Features for Discriminative Localization”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2016, pp. 2921–2929 (cit. on p. 164).

[Zho+22] Xilong Zhou, Miloš Hašan, Valentin Deschaintre, Paul Guerrero, Kalyan Sunkavalli, and Nima Kalantari. “TileGen: Tileable, Controllable Material Generation and Capture”. In: ACM Transactions on Graphics (Proc. SIGGRAPH Asia) (2022) (cit. on pp. 105, 107, 109, 116, 123).

[Zho+23] Xilong Zhou, Miloš Hašan, Valentin Deschaintre, Paul Guerrero, Kalyan Sunkavalli, and Nima Khademi Kalantari. “A Semi-Procedural Convolutional Material Prior”. In: Computer Graphics Forum. Wiley Online Library. 2023 (cit. on p. 105).

[ZK21] Xilong Zhou and Nima Khademi Kalantari. “Adversarial Single-Image SVBRDF Estimation with Hybrid Training”. In: Computer Graphics Forum. Vol. 40. 2. Wiley Online Library. 2021, pp. 315–325 (cit. on pp. 123, 125, 127, 137, 138, 141, 152, 153).

[Zho+17] Yang Zhou, Huajie Shi, Dani Lischinski, Minglun Gong, Johannes Kopf, and Hui Huang. “Analysis and Controlled Synthesis of Inhomogeneous Textures”. In: Computer Graphics Forum. Vol. 36. 2. 2017, pp. 199–212 (cit. on p. 74).

[Zho+18] Yang Zhou, Zhen Zhu, Xiang Bai, Dani Lischinski, Daniel Cohen-Or, and Hui Huang. “Non-Stationary Texture Synthesis by Adversarial Expansion”. In: ACM Transactions on Graphics (TOG) 37.4 (July 2018) (cit. on pp. 34, 38, 56, 72, 74, 76, 77, 79, 84, 85, 87, 90, 94, 104).

[Zho+19] Yizhou Zhou, Xiaoyan Sun, Zheng-Jun Zha, and Wenjun Zeng. “Context-Reinforced Semantic Segmentation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4046–4055 (cit. on p. 41).

[Zhu+17a] Jun Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. “Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2017-Octob (Mar. 2017), pp. 2242–2251 (cit. on pp. 78, 79, 83, 94).

[Zhu+17b] Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A Efros, Oliver Wang, and Eli Shechtman. “Toward Multimodal Image-to-Image Translation”. In: Advances in Neural Information Processing Systems. 2017, pp. 465–476 (cit. on p. 37).

[Zhu+21] Wei Zhu, Xian Guo, Dai Owaki, Kyo Kutsuzawa, and Mitsuhiro Hayashibe. “A Survey of Sim-to-real Transfer Techniques Applied to Reinforcement Learning for Bioinspired Robots”. In: IEEE Transactions on Neural Networks and Learning Systems (2021) (cit. on pp. 156, 212).

_______________________________

1 Additional publication details and results are included on the project website.