<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE book PUBLIC "-//NLM//DTD BITS Book Interchange DTD v2.0 20151225//EN" "https://jats.nlm.nih.gov/extensions/bits/2.1/BITS-book2-1.dtd">
<book book-type="chapter" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xi="http://www.w3.org/2001/XInclude" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="2.0" xml:lang="en">
<book-meta>
<book-id book-id-type="publisher">Universidad Rey Juan Carlos</book-id>
<book-id book-id-type="doi">10.33732/TD-43</book-id>
<subj-group>
<subject>Tesis doctoral</subject>
</subj-group>
<book-title-group>
<book-title><target target-type="page" id="pges_i"/><target target-type="page" id="pges_ii"/><target target-type="page" id="pges_iii"/>Neural networks for digital materials and radiance encoding</book-title>
</book-title-group>
<contrib-group>
<contrib contrib-type="author">
<name name-style="western">
    <surname> Rodr&#x00ED;guez-Pardo Jim&#x00E9;nez</surname>
    <given-names>Carlos Miguel</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name name-style="western">
    <surname>Garc&#x00E9;s Garc&#x00ED;a</surname>
<given-names>Elena</given-names>
</name>
<role>Supervisor</role>
</contrib>
</contrib-group>
<pub-date>
<year>2026</year>
</pub-date>
<publisher>
<publisher-name>Doctoral Program in Information and Communication Technologies.</publisher-name>
<publisher-name>International Doctoral School</publisher-name>
<publisher-loc>Madrid, Spain</publisher-loc>
</publisher>
<permissions>
<copyright-statement>Derechos de autor 2023 los autores</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>los autores</copyright-holder>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-sa/4.0/" xml:lang="en">
<license-p>CC BY-SA 4.0 attribution-noncommercial-sharealike 4.0 international</license-p>
</license>
</permissions>
</book-meta>
<front-matter>
<front-matter-part id="fm01" book-part-type="copyright-page">
<book-part-meta>
<book-part-id book-part-id-type="publisher">URJC</book-part-id>
<title-group>
<title>Copyright</title>
</title-group>
</book-part-meta>
<named-book-part-body>
<p><target target-type="page" id="pges_iv"/><target target-type="page" id="pges_v"/><target target-type="page" id="pges_vi"/>This monograph has been published in collaboration with the Office of Free Knowledge and Culture and the International Doctoral School for having received the extraordinary award at Rey Juan Carlos University.</p>
<p>How to cite</p>
    <p>Rodr&#x00ED;guez-Pardo Jim&#x00E9;nez, C. M. (2023). <italic>Neural networks for digital materials and radiance encoding</italic>. Editorial Academia Abierta. Retrieved from <uri xlink:href="https://monografias.urjc.es/index.php/omp-urjc/catalog/book/43">https://monografias.urjc.es/index.php/omp-urjc/catalog/book/43</uri></p>
<p>Deposit</p>
<p>URJC open archive (digital BURJC)</p>
<p>DOI: <uri xlink:href="https://doi.org/10.33732/TD-43">https://doi.org/10.33732/TD-43</uri></p>
<p>&#x00A9; 2023. Carlos Miguel Rodr&#x00ED;guez-Pardo Jim&#x00E9;nez</p>
<p>&#x00A9; Of the texts: Carlos Miguel Rodr&#x00ED;guez-Pardo Jim&#x00E9;nez</p>
<p>&#x00A9; Of the images: their authors</p>
<p>&#x00A9; Cover image: JSCreative-LabAI_Arts, <uri xlink:href="https://pixabay.com/es/illustrations/ai-generado-cerebro-ondas-cerebrales-8620473/">https://pixabay.com/es/illustrations/ai-generado-cerebro-ondas-cerebrales-8620473/</uri></p>
<p>Editorial</p>
<p>Academia Abierta. Servicio de publicaciones de la URJC</p>
<p>C/Tulip&#x00E1;n s/n</p>
<p>28933 M&#x00F3;stoles</p>
<p><uri xlink:href="https://monografias.urjc.es/index.php/omp-urjc">https://monografias.urjc.es/index.php/omp-urjc</uri></p>
<p>Interior design and cover</p>
<p>Glaux Publicaciones Acad&#x00E9;micas, SLU</p>
<p>Editorial contact</p>
<p>Laura de la Cruz / Tom&#x00E1;s Zarza</p>
<p><email>servicio.publicaciones@urjc.es</email></p>
<p>Some rights reserved</p>
<p>This document is distributed under the Creative Commons Attribution-ShareAlike 4.0 International licence, available at <uri xlink:href="https://creativecommons.org/licenses/by-sa-4.0/">https://creativecommons.org/licenses/by-sa-4.0/</uri></p>
</named-book-part-body>
</front-matter-part>
<dedication id="fm02">
<book-part-meta>
<title-group>
<title>Dedication</title>
</title-group>
</book-part-meta>
<named-book-part-body>
<p><target target-type="page" id="pges_vii"/><italic>That he not busy being born Is busy dying.</italic></p>
<p>&#x2014; <bold>Bob Dylan</bold></p>
<p>It&#x2019;s Alright, Ma (I&#x2019;m Only Bleeding)</p>
</named-book-part-body>
</dedication>
<front-matter-part id="fm03">
<book-part-meta>
<book-part-id book-part-id-type="publisher">URJC</book-part-id>
<title-group>
<title><target target-type="page" id="pges_viii"/><target target-type="page" id="pges_ix"/>Abstract</title>
</title-group>
</book-part-meta>
<named-book-part-body>
<p>Realistic virtual scenes are becoming increasingly prevalent in our society, with a wide range of applications in areas such as manufacturing, architecture, fashion design, and entertainment, including movies, video games, and augmented and virtual reality. Generating realistic images of such scenes requires highly accurate illumination, geometry, and material models, which can be time-consuming and challenging to obtain. Traditionally, such models have often been created manually by skilled artists, but this process can be prohibitively time-consuming and costly. Alternatively, real-world examples can be captured, but this approach presents additional challenges in terms of accuracy and scalability. Moreover, while realism and accuracy are crucial in such processes, rendering efficiency is also a key requirement, so that lifelike images can be generated with the speed required in many real-world applications. One of the most significant challenges in this regard is the acquisition and representation of materials, which are a critical component of our visual world and, by extension, of virtual representations of it. However, existing approaches for material acquisition and representation are limited in terms of efficiency and accuracy, which limits their real-world impact. To address these challenges, data-driven approaches that leverage machine learning may provide viable solutions. Nevertheless, designing and training machine learning models that meet all these competing requirements remains a challenging task, requiring careful consideration of trade-offs between quality and efficiency.</p>
<p><target target-type="page" id="pges_x"/>In this thesis, we propose novel learning-based solutions to address several key challenges in physically-based rendering and material digitization. Our approach leverages various forms of neural networks to introduce innovative algorithms for radiance encoding, digital material generation, edition, and estimation. First, we present a visual attribute transfer framework for digital materials that can effectively generalize to new illumination conditions and geometric distortions. We showcase a use-case of this method for high-resolution material acquisition using a custom device. Additionally, we propose a generative model capable of synthesizing tileable textures from a single input image, which helps improve the quality of material rendering. Building upon recent work in neural fields, we also introduce a material representation that accurately encodes material reflectance while offering powerful editing and propagation capabilities. In addition to reflectance, we present a novel method for global illumination encoding that leverages carefully designed generative models to achieve significantly faster sampling than previous work. Finally, we propose two innovative methods for low-cost material digitization. With flatbed scanners as our capture device, we present a generative model that can provide high-resolution material reflectance estimations using a single image as input, while introducing an uncertainty quantification algorithm that increases its reliability and efficiency. Additionally, we present a novel method for digitizing fabric mechanical properties using depth images as input, which we extend with a perceptually-validated drape similarity metric. Overall, the contributions of this thesis represent significant advances in the fields of radiance encoding and digital material acquisition and edition, enhancing the quality, scalability, and efficiency of physically-based rendering pipelines.</p>
</named-book-part-body>
</front-matter-part>
<ack id="ack1">
<title><target target-type="page" id="pges_xi"/>Acknowledgements</title>
<p>This thesis is the result of several years of relentless hard work. However, none of this would have been possible without the support of an outstanding group of people, to whom I owe a tremendous debt of gratitude. First and foremost, to my advisor Elena Garc&#x00E9;s, for the opportunity to embark on this journey, for everything I have learned from her, for making sure I had all the resources I needed to get my work done, and for the countless and tireless hours she spent helping me improve as a researcher. I would also like to express my heartfelt appreciation to Dan Casas and Jorge Lopez-Moreno, who have been incredible mentors, always ready to offer guidance and support, and from whom I have learned so much over the years.</p>
<p>I want to express my sincere gratitude to the team at SEDDI for their unwavering support over the years. I feel specially grateful for the opportunity to work closely with David Pascual, Sergio Ruiz, and Sofia Dom&#x00ED;nguez, who not only helped me professionally but also made our time together more fun and enjoyable (with great <italic>padel</italic> matches and even better inside jokes, and silly names for our AI model iterations). They have been the best teammates I could have ever asked for. In no particular order, I would also like to thank Henar Dom&#x00ED;nguez, Sergio Suja, Javier Fabre, Loreto P&#x00E9;rez, Daniel Rodr&#x00ED;guez, and Luis Romero, who have also made this journey much more fun and have positively contributed in many ways to my work.</p>
<p>I would like to express my deepest appreciation to SEDDI as a company, which I joined shortly after it was founded. Their financial and moral support <target target-type="page" id="pges_xii"/>has been crucial to the successful completion of this thesis, and I am grateful for their unwavering assistance throughout this journey. Thanks to their generosity, I was able to access all the resources I needed, and they allowed me to publicly share my work through publications. I believe that SEDDI&#x2019;s future is bright, and I am excited to witness their future accomplishments.</p>
<p>I would also like to thank my friends and family, for their support and patience over the past years; my students, for their trust; and the people at the MSLab at URJC and the GIAA at UC3M who have helped me throughout these years.</p>
<p>I want to acknowledge the invaluable contributions of the reviewers who have selflessly dedicated their time and effort to provide constructive feedback on our papers. Their rigorous reviews have been instrumental in improving the quality of our research. Moreover, I am deeply grateful to the open-source community and the authors of countless open-access papers that have inspired and shaped many of the ideas we have developed in this work. I firmly believe that open communication and collaboration are critical for advancing the fields of machine learning and computer vision. Thus, I sincerely hope that the research communities continue to publish their work, preferably with open-source implementations and data, to facilitate knowledge-sharing and foster innovation.</p>
<p>Last but not least, I am grateful for the public investment in research and innovation, as well as the Spanish public education system, without which this thesis would not have been possible. Looking forward, I sincerely hope that as a society, we can redouble our efforts to defend our public democratic institutions, which are crucial for building a fair, free, sustainable, and equitable world. It is imperative that we prioritize investing in public education and research to foster innovation, create new opportunities, and address the complex challenges we face today. By doing so, we can build a brighter future for ourselves and future generations.</p>
</ack>
<toc id="fm04" content-type="toc">
<toc-title-group>
<title><target target-type="page" id="pges_xiii"/>Contents</title>
</toc-title-group>
<toc-entry><title>Abstract</title> <nav-pointer rid="fm03">ix</nav-pointer></toc-entry>
    <toc-entry><title>Acknowledgements</title> <nav-pointer rid="ack1">xi</nav-pointer></toc-entry>
<toc-entry content-type="chapter"><label>1</label> <title>Introduction</title> <nav-pointer rid="c1">1</nav-pointer>
<toc-entry content-type="section"><label>&#x2022; 1.1</label> <title>Theoretical Background</title> <nav-pointer rid="c1-s1">3</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 1.1.1</label> <title>Physically-Based Rendering</title> <nav-pointer rid="c1-s1.s1">3</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 1.1.2</label> <title>Geometry</title> <nav-pointer rid="c1-s1.s2">5</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 1.1.3</label> <title>Materials</title> <nav-pointer rid="c1-s1.s3">8</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 1.1.4</label> <title>Illumination</title> <nav-pointer rid="c1-s1.s4">12</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 1.2</label> <title>Open Problems</title> <nav-pointer rid="c1-s2">14</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 1.3</label> <title>Goals &#x0026; Contributions</title> <nav-pointer rid="c1-s3">19</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 1.3.1</label> <title>Contributions</title> <nav-pointer rid="c1-s3.s1">20</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 1.4</label> <title>Measurable Outcomes</title> <nav-pointer rid="c1-s4">23</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 1.4.1</label> <title>Publications &#x0026; Patents</title> <nav-pointer rid="c1-s4.s1">23</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 1.4.2</label> <title>Awards &#x0026; Other Merits</title> <nav-pointer rid="c1-s4.s2">24</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 1.4.3</label> <title>Student Supervision</title> <nav-pointer rid="c1-s4.s3">25</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 1.4.4</label> <title>Teaching, Talks, and Seminars</title> <nav-pointer rid="c1-s4.s4">25</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 1.4.5</label> <title>Reviewing Activity</title> <nav-pointer rid="c1-s4.s5">26</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 1.4.6</label> <title>Industrial Research Projects and Commercial Products</title> <nav-pointer rid="c1-s4.s6">26</nav-pointer></toc-entry>
</toc-entry>
</toc-entry>
<toc-entry content-type="chapter"><label>2</label> <title>Visual Attribute Transfer with Photometry-guided Neural Networks</title> <nav-pointer rid="c2">27</nav-pointer>
<toc-entry content-type="section"><label>&#x2022; 2.1</label> <title>Introduction</title> <nav-pointer rid="c2-s1">28</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.2</label> <title>Related Work</title> <nav-pointer rid="c2-s2">30</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 2.2.1</label> <title>Visual Attribute Transfer</title> <nav-pointer rid="c2-s2.s1">30</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.2.2</label> <title>Textured Materials</title> <nav-pointer rid="c2-s2.s2">31</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.2.3</label> <title>SVBRDF Estimation</title> <nav-pointer rid="c2-s2.s3">32</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.3</label> <title>Problem Formulation</title> <nav-pointer rid="c2-s3">33</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.4</label> <title>Learning Framework</title> <nav-pointer rid="c2-s4">34</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 2.4.1</label> <title>Patch-based Training</title> <nav-pointer rid="c2-s4.s1">34</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.4.2</label> <title>Network Design</title> <nav-pointer rid="c2-s4.s2">35</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.4.3</label> <title>Implementation details</title> <nav-pointer rid="c2-s4.s3">37</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.5</label> <title>Dataset and Metrics</title> <nav-pointer rid="c2-s5">38</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 2.5.1</label> <title>Dual-Resolution Captured Data</title> <nav-pointer rid="c2-s5.s1">38</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.5.2</label> <title>Attribute-specific Metrics for Evaluation</title> <nav-pointer rid="c2-s5.s2">39</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.6</label> <title>Evaluation</title> <nav-pointer rid="c2-s6">39</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 2.6.1</label> <title>Invariance to Input Illuminations and Distortions</title> <nav-pointer rid="c2-s6.s1">39</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.6.2</label> <title>Equivariance to Affine Transforms</title> <nav-pointer rid="c2-s6.s2">41</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label><target target-type="page" id="pges_xiv"/>&#x2022; 2.7</label> <title>Results and Comparisons</title> <nav-pointer rid="c2-s7">43</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 2.7.1</label> <title>Interactive Stylizations</title> <nav-pointer rid="c2-s7.s1">43</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.7.2</label> <title>Creation of Large Scale Digital Material Assets</title> <nav-pointer rid="c2-s7.s2">45</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.7.3</label> <title>Generalization to Similar Materials</title> <nav-pointer rid="c2-s7.s3">49</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.7.4</label> <title>Limitations</title> <nav-pointer rid="c2-s7.s4">49</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.8</label> <title>Conclusions and Future Work</title> <nav-pointer rid="c2-s8">49</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.A</label> <title>Additional Implementation Details</title> <nav-pointer rid="c2-sA">52</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.B</label> <title>Additional Comparisons</title> <nav-pointer rid="c2-sB">54</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 2.C</label> <title>Extension to SVBSDF Propagation</title> <nav-pointer rid="c2-sC">60</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 2.C.1</label> <title>Mesoscale Propagation</title> <nav-pointer rid="c2-sC.s1">61</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.C.2</label> <title>Implementation Details</title> <nav-pointer rid="c2-sC.s2">63</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 2.C.3</label> <title>Results</title> <nav-pointer rid="c2-sC.s3">66</nav-pointer></toc-entry>
</toc-entry>
</toc-entry>
<toc-entry content-type="chapter"><label>3</label> <title>A Single Image Generative Model for Tileable Texture Synthesis</title> <nav-pointer rid="c3">69</nav-pointer>
<toc-entry content-type="section"><label>&#x2022; 3.1</label> <title>Introduction</title> <nav-pointer rid="c3-s1">70</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.2</label> <title>Related work</title> <nav-pointer rid="c3-s2">71</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 3.2.1</label> <title>Texture Synthesis</title> <nav-pointer rid="c3-s2.s1">71</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 3.2.2</label> <title>Texture Tileability</title> <nav-pointer rid="c3-s2.s2">73</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 3.2.3</label> <title>Deep Internal Learning</title> <nav-pointer rid="c3-s2.s3">73</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.3</label> <title>Overview</title> <nav-pointer rid="c3-s3">74</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.4</label> <title>Self-Supervised Texture Synthesis with Adversarial Expansion</title> <nav-pointer rid="c3-s4">75</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 3.4.1</label> <title>Network Architecture</title> <nav-pointer rid="c3-s4.s1">76</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.5</label> <title>Tileable Texture Stack Sampling</title> <nav-pointer rid="c3-s5">78</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 3.5.1</label> <title>Latent Space Tiling</title> <nav-pointer rid="c3-s5.s1">78</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 3.5.2</label> <title>Discriminator-guided Sampling</title> <nav-pointer rid="c3-s5.s2">79</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.6</label> <title>Implementation Details</title> <nav-pointer rid="c3-s6">81</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.7</label> <title>Experiments</title> <nav-pointer rid="c3-s7">83</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 3.7.1</label> <title>Generator <inline-formula><mml:math id="M1" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> Design</title> <nav-pointer rid="c3-s7.s1">83</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 3.7.2</label> <title>Inter-map Consistency</title> <nav-pointer rid="c3-s7.s2">84</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 3.7.3</label> <title>Loss Function Ablation Study</title> <nav-pointer rid="c3-s7.s3">85</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.8</label> <title>Results and Comparisons</title> <nav-pointer rid="c3-s8">86</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.9</label> <title>Conclusions and Future Work</title> <nav-pointer rid="c3-s9">89</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.A</label> <title>Additional Details</title> <nav-pointer rid="c3.sA">92</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 3.B</label> <title>Additional Comparisons, Results and Ablations</title> <nav-pointer rid="c3.sB">93</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="chapter"><label>4</label> <title>Neural Fields for BTF Encoding and Transfer</title> <nav-pointer rid="c4">99</nav-pointer>
<toc-entry content-type="section"><label>&#x2022; 4.1</label> <title>Introduction</title> <nav-pointer rid="c4-s1">100</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 4.2</label> <title>Related Work</title> <nav-pointer rid="c4-s2">101</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 4.3</label> <title>Method</title> <nav-pointer rid="c4-s3">104</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 4.3.1</label> <title>Inference</title> <nav-pointer rid="c4-s3.s1">104</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 4.3.2</label> <title>Training</title> <nav-pointer rid="c4-s3.s2">105</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 4.4</label> <title>Implementation details</title> <nav-pointer rid="c4-s4">107</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 4.5</label> <title>Evaluation</title> <nav-pointer rid="c4-s5">109</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 4.5.1</label> <title>Qualitative Analysis</title> <nav-pointer rid="c4-s5.s1">109</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 4.5.2</label> <title>Compression comparisons with previous work</title> <nav-pointer rid="c4-s5.s2">110</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 4.5.3</label> <title>Limitations</title> <nav-pointer rid="c4-s5.s3">111</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 4.6</label> <title>Applications</title> <nav-pointer rid="c4-s6">111</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 4.6.1</label> <title>Reflectance Propagation and Tileable BTFs</title> <nav-pointer rid="c4-s6.s1">111</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 4.6.2</label> <title>Structural Material Edition</title> <nav-pointer rid="c4-s6.s2">113</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 4.6.3</label> <title>Multi Resolution Neural Materials</title> <nav-pointer rid="c4-s6.s3">113</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 4.7</label> <title>Conclusions</title> <nav-pointer rid="c4-s7">114</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 4.A</label> <title>Model Implementation Details</title> <nav-pointer rid="c4-sA">117</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="chapter"><label>5</label> <title>Uncertainty-Aware Neural Networks for High Resolution Digitization</title> <nav-pointer rid="c5">121</nav-pointer>
<toc-entry content-type="section"><label>&#x2022; 5.1</label> <title>Introduction</title> <nav-pointer rid="c5-s1">123</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 5.2</label> <title>Related Work</title> <nav-pointer rid="c5-s2">125</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 5.2.1</label> <title>Lightweight Material Capture</title> <nav-pointer rid="c5-s2.s1">125</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.2.2</label> <title>Uncertainty Quantification in Deep Learning</title> <nav-pointer rid="c5-s2.s2">126</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 5.3</label> <title>Method</title> <nav-pointer rid="c5-s3">127</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 5.3.1</label> <title>Network Design</title> <nav-pointer rid="c5-s3.s1">127</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.3.2</label> <title>Loss Function</title> <nav-pointer rid="c5-s3.s2">128</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.3.3</label> <title>Data Augmentation</title> <nav-pointer rid="c5-s3.s3">129</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 5.4</label> <title>Uncertainty Quantification</title> <nav-pointer rid="c5-s4">129</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 5.5</label> <title>Experimental Results</title> <nav-pointer rid="c5-s5">131</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 5.5.1</label> <title>Dataset</title> <nav-pointer rid="c5-s5.s1">131</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.5.2</label> <title>Metrics</title> <nav-pointer rid="c5-s5.s2">131</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 5.6</label> <title>Evaluation</title> <nav-pointer rid="c5-s6">132</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 5.6.1</label> <title>Ablation Study</title> <nav-pointer rid="c5-s6.s1">132</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.6.2</label> <title>Qualitative Analysis</title> <nav-pointer rid="c5-s6.s2">132</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.6.3</label> <title>Uncertainty Evaluation</title> <nav-pointer rid="c5-s6.s3">135</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.6.4</label> <title>Active Learning</title> <nav-pointer rid="c5-s6.s4">135</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.6.5</label> <title>Comparisons with Previous Work</title> <nav-pointer rid="c5-s6.s5">138</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.6.6</label> <title>Limitations</title> <nav-pointer rid="c5-s6.s6">140</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 5.7</label> <title>Conclusions</title> <nav-pointer rid="c5-s7">140</nav-pointer></toc-entry>
<toc-entry content-type="section"><label><target target-type="page" id="pges_xvi"/>&#x2022; 5.A</label> <title>Additional Implementation Details</title> <nav-pointer rid="c5.sA">140</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 5.A.1</label> <title>Model Design</title> <nav-pointer rid="c5.sA.s1">140</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.A.2</label> <title>Model Training</title> <nav-pointer rid="c5.sA.s2">141</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.A.3</label> <title>Artifact Detection</title> <nav-pointer rid="c5.sA.s3">142</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.A.4</label> <title>Capture Details</title> <nav-pointer rid="c5.sA.s4">143</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 5.A.5</label> <title>Comparisons with Previous Work</title> <nav-pointer rid="c5.sA.s5">143</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 5.B</label> <title>Dataset Analysis</title> <nav-pointer rid="c5.sB">146</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 5.C</label> <title>Additional Details And Results</title> <nav-pointer rid="c5.sC">148</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="chapter"><label>6</label> <title>Casual Capture of Fabric Mechanics</title> <nav-pointer rid="c6">155</nav-pointer>
<toc-entry content-type="section"><label>&#x2022; 6.1</label> <title>Introduction</title> <nav-pointer rid="c6-s1">156</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.2</label> <title>Related Work</title> <nav-pointer rid="c6-s2">158</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 6.2.1</label> <title>Parameter Estimation Methods</title> <nav-pointer rid="c6-s2.s1">158</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 6.2.2</label> <title>Pre-Trained Models and Transfer Learning</title> <nav-pointer rid="c6-s2.s2">160</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 6.2.3</label> <title>Similarity Metrics</title> <nav-pointer rid="c6-s2.s3">160</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.3</label> <title>Overview</title> <nav-pointer rid="c6-s3">161</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.4</label> <title>Datasets</title> <nav-pointer rid="c6-s4">162</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 6.4.1</label> <title>Data Augmentation</title> <nav-pointer rid="c6-s4.s1">163</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.5</label> <title>Fabric Mechanics from Depth Images</title> <nav-pointer rid="c6-s5">165</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 6.5.1</label> <title>Neural Network Architecture</title> <nav-pointer rid="c6-s5.s1">165</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 6.5.2</label> <title>Quantitative Evaluation</title> <nav-pointer rid="c6-s5.s2">165</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.6</label> <title>A Similarity Metric for Drape</title> <nav-pointer rid="c6-s6">169</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 6.6.1</label> <title>Image-based Similarity of Drape</title> <nav-pointer rid="c6-s6.s1">169</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.7</label> <title>Evaluation</title> <nav-pointer rid="c6-s7">171</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 6.7.1</label> <title>Human Judgment Perceptual Similarity of Drape</title> <nav-pointer rid="c6-s7.s1">172</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 6.7.2</label> <title>Image-based vs. Human Perceptual Similarity</title> <nav-pointer rid="c6-s7.s2">172</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 6.7.3</label> <title>Qualitative Results</title> <nav-pointer rid="c6-s7.s3">174</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.8</label> <title>Conclusions</title> <nav-pointer rid="c6-s8">174</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.A</label> <title>Additional Implementation Details</title> <nav-pointer rid="c6-sA">177</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.B</label> <title>Additional Results</title> <nav-pointer rid="c6-sB">177</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 6.C</label> <title>Ablation Study: Similarity Metric Parameters</title> <nav-pointer rid="c6-sC">181</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="chapter"><label>7</label> <title>Neural Models for Global Illumination</title> <nav-pointer rid="c7">187</nav-pointer>
<toc-entry content-type="section"><label>&#x2022; 7.1</label> <title>Introduction</title> <nav-pointer rid="c7-s1">188</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 7.2</label> <title>Related Work</title> <nav-pointer rid="c7-s2">189</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 7.2.1</label> <title>Neural Sampling and Representations</title> <nav-pointer rid="c7-s2.s1">190</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 7.3</label> <title>Overview</title> <nav-pointer rid="c7-s3">192</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 7.3.1</label> <title>Map Pre-processing</title> <nav-pointer rid="c7-s3.s1">193</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label><target target-type="page" id="pges_xvii"/>&#x2022; 7.3.2</label> <title>Computing Ground Truth Tabular PDFs</title> <nav-pointer rid="c7-s3.s2">193</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 7.4</label> <title>Learned Sampling and PDF evaluation</title> <nav-pointer rid="c7-s4">194</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 7.4.1</label> <title>Formalization</title> <nav-pointer rid="c7-s4.s1">194</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 7.4.2</label> <title>Design of <inline-formula><mml:math id="M2" display='block'><mml:mo mathvariant="script">F</mml:mo></mml:math></inline-formula></title> <nav-pointer rid="c7-s4.s2">195</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 7.4.3</label> <title>Training <inline-formula><mml:math id="M3" display='block'><mml:mo mathvariant="script">F</mml:mo></mml:math></inline-formula></title> <nav-pointer rid="c7-s4.s3">196</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 7.5</label> <title>Environment Map Compression</title> <nav-pointer rid="c7-s5">197</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 7.6</label> <title>Implementation Details</title> <nav-pointer rid="c7-s6">197</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 7.7</label> <title>Dataset</title> <nav-pointer rid="c7-s7">199</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; 7.8</label> <title>Evaluation</title> <nav-pointer rid="c7-s8">199</nav-pointer>
<toc-entry content-type="subsection"><label>&#x2022; 7.8.1</label> <title>PDF Fit Accuracy</title> <nav-pointer rid="c7-s8.s1">200</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 7.8.2</label> <title>Environment Map Reconstruction</title> <nav-pointer rid="c7-s8.s2">202</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 7.8.3</label> <title>Computational Cost</title> <nav-pointer rid="c7-s8.s3">204</nav-pointer></toc-entry>
<toc-entry content-type="subsection"><label>&#x2022; 7.8.4</label> <title>Rendering Comparisons</title> <nav-pointer rid="c7-s8.s4">205</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="section"><label>&#x2022; 7.9</label> <title>Conclusions</title> <nav-pointer rid="c7-s9">206</nav-pointer></toc-entry>
</toc-entry>
<toc-entry content-type="chapter"><label>8</label> <title>Conclusions</title> <nav-pointer rid="c8">209</nav-pointer></toc-entry>
<toc-entry content-type="bibliography"><title>Bibliography</title> <nav-pointer rid="c9">219</nav-pointer></toc-entry>
<toc-entry content-type="chapter"><title>A Resumen</title> <nav-pointer rid="c10-app1">269</nav-pointer>
<toc-entry content-type="section"><label>&#x2022; A.1</label> <title>Antecedentes</title> <nav-pointer rid="c10-app1.s1">270</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; A.2</label> <title>Objetivos</title> <nav-pointer rid="c10-app1.s2">274</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; A.3</label> <title>Metodolog&#x00ED;a</title> <nav-pointer rid="c10-app1.s3">275</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; A.4</label> <title>Resultados</title> <nav-pointer rid="c10-app1-s4">280</nav-pointer></toc-entry>
<toc-entry content-type="section"><label>&#x2022; A.5</label> <title>Conclusiones</title> <nav-pointer rid="c10-app1-s5">283</nav-pointer></toc-entry>
</toc-entry>
</toc>
<toc id="fm05" content-type="figures">
<toc-title-group>
<title><target target-type="page" id="pges_xviii"/><target target-type="page" id="pges_xix"/>Figures</title>
</toc-title-group>
<toc-entry content-type="figure"><label>&#x2022; 1.1</label> <title>Renders of virtual scenes for different use cases.</title> <nav-pointer rid="fig-1.1">1</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.2</label> <title>Images of different real-world materials and textures.</title> <nav-pointer rid="fig-1.2">2</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.3</label> <title>Example of light transport and material interactions.</title> <nav-pointer rid="fig-1.3">4</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.4</label> <title>Example of light transport, connecting light sources to pixels.</title> <nav-pointer rid="fig-1.4">6</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.5</label> <title>2D depiction of a physically-principled BSDF theorical model.</title> <nav-pointer rid="fig-1.5">7</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.6</label> <title>Example of spatially varying illumination encoding with Spherical Harmonics.</title> <nav-pointer rid="fig-1.6">12</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.7</label> <title>Schemes and photos of SEDDI&#x2019;s optical device</title> <nav-pointer rid="fig-1.7">15</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.8</label> <title>Example of a non-tileable and a tileable material.</title> <nav-pointer rid="fig-1.8">16</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.9</label> <title>Renders of real and simulated garments.</title> <nav-pointer rid="fig-1.9">18</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.10</label> <title>Overview of the structure of this thesis.</title> <nav-pointer rid="fig-1.10">19</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 1.11</label> <title>Lightweight material digitization enabled by the methods presented in this thesis.</title> <nav-pointer rid="fig-1.11">22</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.1</label> <title>Illustration of the capabilities of our attribute transfer framework.</title> <nav-pointer rid="fig-2.1">27</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.2</label> <title>Teaser Figure for <italic>photometricNet</italic></title> <nav-pointer rid="fig-2.2">28</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.3</label> <title>Overview of <italic>photometricNet</italic></title> <nav-pointer rid="fig-2.3">33</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.4</label> <title>Illustration of our proposed color augmentation policy</title> <nav-pointer rid="fig-2.4">34</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.5</label> <title>An overview of our evaluation dataset for <italic>PhotometricNet</italic></title> <nav-pointer rid="fig-2.5">36</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.6</label> <title>Qualitative results of our models under different datasets and data augmentation configurations.</title> <nav-pointer rid="fig-2.6">38</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.7</label> <title>Ablation study on the influence of the size of the photometric dataset on the quality of the propagations.</title> <nav-pointer rid="fig-2.7">41</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.8</label> <title>Output of our model under different distorted inputs.</title> <nav-pointer rid="fig-2.8">42</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.9</label> <title>Quantitative evaluation of the robustness of our networks with respect to different transformations.</title> <nav-pointer rid="fig-2.9">43</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.10</label> <title>Comparisons of our method with previous work on stylizations.</title> <nav-pointer rid="fig-2.10">44</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.11</label> <title>Comparisons of our method with previous work on image analogies.</title> <nav-pointer rid="fig-2.11">46</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.12</label> <title>Results of our framework for interactive material-aware visual attribute transfer.</title> <nav-pointer rid="fig-2.12">46</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.13</label> <title>Results of our framework for material capture using a smartphone</title> <nav-pointer rid="fig-2.13">47</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.14</label> <title>Qualitative comparison of photometricNet with previous work on guided attribute transfer</title> <nav-pointer rid="fig-2.14">47</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.15</label> <title>Results of our method on real BTF captured data.</title> <nav-pointer rid="fig-2.15">48</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.16</label> <title>Generalization capabilities of photometricNet when evaluated on materials similar to those in their training dataset.</title> <nav-pointer rid="fig-2.16">50</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label><target target-type="page" id="pges_xx"/>&#x2022; 2.17</label> <title>Failure cases of photometricNet.</title> <nav-pointer rid="fig-2.17">51</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.18</label> <title>An overview of our deep learning architecture</title> <nav-pointer rid="fig-2.18">53</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.19</label> <title>Impact of the depth of the model on its accuracy</title> <nav-pointer rid="fig-2.19">53</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.20</label> <title>Impact of the choice of the loss function used to train the networks in the estimated maps at macroscale</title> <nav-pointer rid="fig-2.20">55</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.21</label> <title>Relative position of the 27 directional light sources used to capture the datasets to train photometricNet.</title> <nav-pointer rid="fig-2.21">56</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.22</label> <title>Additional comparison of photometricNet with related image analogies methods.</title> <nav-pointer rid="fig-2.22">58</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.23</label> <title>Additional comparisons of photometricNet with related image stylization methods.</title> <nav-pointer rid="fig-2.23">59</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.24</label> <title>A comparison of photometricNet with material digitization methods.</title> <nav-pointer rid="fig-2.24">60</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.25</label> <title>Renders of SVBSDF maps transferred by our method.</title> <nav-pointer rid="fig-2.25">60</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.26</label> <title>Virtual scene rendered using many materials generated by our method.</title> <nav-pointer rid="fig-2.26">61</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.27</label> <title>Comparison of the extended SVBSDF method with the baseline photometricNet.</title> <nav-pointer rid="fig-2.27">64</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.28</label> <title>Microscale SVBSDF maps.</title> <nav-pointer rid="fig-2.28">65</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.29</label> <title>A full diagram of the model architecture of our extended photometricNet.</title> <nav-pointer rid="fig-2.29">67</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 2.30</label> <title>Ablation study of our proposed improvements with respect to photometricNet.</title> <nav-pointer rid="fig-2.30">68</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.1</label> <title>Results obtained with SeamlessGAN for tileable texture synthesis.</title> <nav-pointer rid="fig-3.1">69</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.2</label> <title>Illustration of the capabilities of SeamlessGAN.</title> <nav-pointer rid="fig-3.2">72</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.3</label> <title>Overview of SeamlessGAN</title> <nav-pointer rid="fig-3.3">74</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.4</label> <title>Training framework of SeamlessGAN.</title> <nav-pointer rid="fig-3.4">76</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.5</label> <title>Impact of the latent field level on the quality of the outputs of SeamlessGAN.</title> <nav-pointer rid="fig-3.5">78</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.6</label> <title>Outputs of the discriminator of SeamlessGAN for textures of different qualities.</title> <nav-pointer rid="fig-3.6">79</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.7</label> <title>Examples of SeamlessGAN generating multiple tileable outputs from the same sample</title> <nav-pointer rid="fig-3.7">80</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.8</label> <title>Crops of synthesized textures using SeamlessGAN.</title> <nav-pointer rid="fig-3.8">82</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.9</label> <title>Ablation study on the impact of the loss function on the quality of the synthesized textures.</title> <nav-pointer rid="fig-3.9">85</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.10</label> <title>A comparison of the results of our proposed architectures.</title> <nav-pointer rid="fig-3.10">86</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.11</label> <title>A comparison of our work with <italic>Texture Stationarization.</italic></title> <nav-pointer rid="fig-3.11">88</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.12</label> <title>Qualitative comparison of SeamlessGAN with different texture synthesis methods.</title> <nav-pointer rid="fig-3.12">90</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label><target target-type="page" id="pges_xxi"/>&#x2022; 3.13</label> <title>Examples of failure cases of SeamlessGAN</title> <nav-pointer rid="fig-3.13">91</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.14</label> <title>An ablation study on the influence of the number of layers in each model in the quality of the generated outputs.</title> <nav-pointer rid="fig-3.14">95</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.15</label> <title>More examples on the capabilities of SeamlessGAN on generating multiple tileable textures from a given input.</title> <nav-pointer rid="fig-3.15">96</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 3.16</label> <title>Additional comparison with previous methods.</title> <nav-pointer rid="fig-3.16">97</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.1</label> <title>Renders generated with NeuBTF materials.</title> <nav-pointer rid="fig-4.1">99</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.2</label> <title>An overview of the inference and training processes of NeuBTF</title> <nav-pointer rid="fig-4.2">102</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label><target target-type="page" id="pges_xv"/>&#x2022; 4.3</label> <title>Some tileable neural materials achieved with NeuBTF</title> <nav-pointer rid="fig-4.3">104</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.4</label> <title>Qualitative results of NeuBTF on a variety of BTFs.</title> <nav-pointer rid="fig-4.4">106</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.5</label> <title>A selection of latent channels learned by NeuBTF for a variety of materials.</title> <nav-pointer rid="fig-4.5">108</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.6</label> <title>Surface normals reconstructed from the materials encoded by NeuBTF</title> <nav-pointer rid="fig-4.6">110</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.7</label> <title>A failure case of NeuBTF</title> <nav-pointer rid="fig-4.7">112</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.8</label> <title>Examples of structural editions allowed by NeuBTF on a variety of materials.</title> <nav-pointer rid="fig-4.8">115</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.9</label> <title>Multi resolution material generation using NeuBTF</title> <nav-pointer rid="fig-4.9">116</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.10</label> <title>A diagram of the autoencoder of NeuBTF</title> <nav-pointer rid="fig-4.10">117</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 4.11</label> <title>A diagram of the renderer of NeuBTF</title> <nav-pointer rid="fig-4.11">118</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.1</label> <title>Teaser illustrating our low-cost optical capture pipeline</title> <nav-pointer rid="fig-5.1">121</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.2</label> <title>Illustration of the variance in rendering obtained by materials with different uncertainties</title> <nav-pointer rid="fig-5.2">122</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.3</label> <title>Scanner images vs fitted albedos.</title> <nav-pointer rid="fig-5.3">123</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.4</label> <title>Overview of UMat</title> <nav-pointer rid="fig-5.4">126</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.5</label> <title>Top: input image of a <italic>rib</italic> material with metallic sequins. Bottom: <inline-formula><mml:math id="M4" display='block'><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mtext>BRDF</mml:mtext></mml:msub></mml:math></inline-formula> and per-map uncertainties.</title> <nav-pointer rid="fig-5.5">129</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.6</label> <title>Qualitative results of some configurations of our ablation study</title> <nav-pointer rid="fig-5.6">133</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.7</label> <title>Embeddings of the transformer of our generator for data in the training set.</title> <nav-pointer rid="fig-5.7">133</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.8</label> <title>Correlations between per-map errors and uncertainties.</title> <nav-pointer rid="fig-5.8">136</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.9</label> <title>Illustration and results of our active learning experiment.</title> <nav-pointer rid="fig-5.9">136</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.10</label> <title>Limitations cases of our method.</title> <nav-pointer rid="fig-5.10">139</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.11</label> <title>A full diagram of the generator of UMat.</title> <nav-pointer rid="fig-5.11">144</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.12</label> <title>A full diagram of the discriminator of UMat.</title> <nav-pointer rid="fig-5.12">145</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.13</label> <title>Visualization of the dataset used to train UMat.</title> <nav-pointer rid="fig-5.13">146</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.14</label> <title>Visualization of some Ground Truth SVBRDF of the different families in our test set.</title> <nav-pointer rid="fig-5.14">147</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label><target target-type="page" id="pges_xxii"/>&#x2022; 5.15</label> <title>Additional qualitative results of our uncertainty quantification method</title> <nav-pointer rid="fig-5.15">148</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.16</label> <title>Additional results of our uncertainty metric</title> <nav-pointer rid="fig-5.16">149</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.17</label> <title>Additional results of our active learning experiment.</title> <nav-pointer rid="fig-5.17">150</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 5.18</label> <title>Results of our method on datasets of previous work.</title> <nav-pointer rid="fig-5.18">151</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.1</label> <title>Teaser illustrating the capabilities of our capture setup and similarity metric for fabric mechanics.</title> <nav-pointer rid="fig-6.1">155</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.2</label> <title>Our proposed setup for capturing fabric mechanical parameter</title> <nav-pointer rid="fig-6.2">157</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.3</label> <title>An overview of the main components of our learning-based mechanical parameter capture method.</title> <nav-pointer rid="fig-6.3">160</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.4</label> <title>Spearman correlation matrix between parameters of our synthetic dataset.</title> <nav-pointer rid="fig-6.4">161</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.5</label> <title>Sweep of simulation parameters for hanging and stretch scenes.</title> <nav-pointer rid="fig-6.5">162</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.6</label> <title>Diagram of our training and evaluation pipelines.</title> <nav-pointer rid="fig-6.6">164</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.7</label> <title>Saliency maps of our proposed model, aggregated on target parameter</title> <nav-pointer rid="fig-6.7">169</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.8</label> <title>Human Judgments 2D tSTE embedding of human perceptual judgments about fabric similarity</title> <nav-pointer rid="fig-6.8">171</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.9</label> <title>Correlation between the ordering provided by the Human Judgments and our drape similarity metric.</title> <nav-pointer rid="fig-6.9">173</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.10</label> <title>A comparison between the simulations obtained through the ground truth parameters and those obtained using the predictions of our model.</title> <nav-pointer rid="fig-6.10">175</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.11</label> <title>Search by drape similarity</title> <nav-pointer rid="fig-6.11">176</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.12</label> <title>A failure case of our method.</title> <nav-pointer rid="fig-6.12">176</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.13</label> <title>Relationship between the distance obtained by our user study, our distance metric and the difference in each parameter</title> <nav-pointer rid="fig-6.13">178</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.14</label> <title>Illustration of our simulation-space data augmentation policy</title> <nav-pointer rid="fig-6.14">178</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.15</label> <title>Ranking classification accuracy on different ways of comparing materials.</title> <nav-pointer rid="fig-6.15">179</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.16</label> <title>Correlation between the ordering provided by the Human Judgments and our Drape Similarity Metric.</title> <nav-pointer rid="fig-6.16">180</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.17</label> <title>Correlation between the ordering provided by our similarity metric with the Ground Truth parameters and the estimations.</title> <nav-pointer rid="fig-6.17">181</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.18</label> <title>Correlation between the ordering provided by the tSTE embedding from our user study and our distance metric, computed using the estimated parameters.</title> <nav-pointer rid="fig-6.18">182</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.19</label> <title>Correlation between the ordering provided by the tSTE embedding from our user study and the distance computed on the parameter space, for the model predictions.</title> <nav-pointer rid="fig-6.19">182</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 6.20</label> <title>Demographics of the participants of our perceptual study</title> <nav-pointer rid="fig-6.20">183</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label><target target-type="page" id="pges_xxiii"/>&#x2022; 6.21</label> <title>Ablation study on different factors of our similarity metric.</title> <nav-pointer rid="fig-6.21">183</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.1</label> <title>Renders using analytical illumination methods and our approach.</title> <nav-pointer rid="fig-7.1">187</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.2</label> <title>Rendered images using traditional and our learning based illumination framework.</title> <nav-pointer rid="fig-7.2">188</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.3</label> <title>Overview of NEnv</title> <nav-pointer rid="fig-7.3">191</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.4</label> <title>Environment map and the PDF encoded by NEnv for an unprocessed and a rotated environment map</title> <nav-pointer rid="fig-7.4">192</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.5</label> <title>Quantitative comparisons of the performance of different coupling layers for encoding environment map PDFs.</title> <nav-pointer rid="fig-7.5">200</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.6</label> <title>A qualitative comparison between the type of coupling layer used in our normalizing flows.</title> <nav-pointer rid="fig-7.6">201</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.7</label> <title>A visualization of the impact of the size of <inline-formula><mml:math id="M5" display='block'><mml:mo mathvariant="script">F</mml:mo></mml:math></inline-formula></title> <nav-pointer rid="fig-7.7">202</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.8</label> <title>Environment maps encoded by our method and previous work.</title> <nav-pointer rid="fig-7.8">203</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.9</label> <title>Average seconds per sample for the baseline sampling algorithms and our model.</title> <nav-pointer rid="fig-7.9">204</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 7.10</label> <title>Render comparisons of different configurations of our work with the ground truth environment map and previous work.</title> <nav-pointer rid="fig-7.10">207</nav-pointer></toc-entry>
<toc-entry content-type="figure"><label>&#x2022; 8.1</label> <title>Renders with materials generated by some of the methods presented in this thesis.</title> <nav-pointer rid="fig-8.1">210</nav-pointer></toc-entry>
</toc>
<toc id="fm06" content-type="tables">
<toc-title-group>
<title><target target-type="page" id="pges_xxiv"/><target target-type="page" id="pges_xxv"/>Tables</title>
</toc-title-group>
<toc-entry content-type="table"><label>&#x2022; 2.1</label> <title>Quantitative comparison of photometricNet with previous work.</title> <nav-pointer rid="c2-tab1">48</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 2.2</label> <title>Quantitative comparison between different number of blocks of layers in our model.</title> <nav-pointer rid="c2-tab2">52</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 3.1</label> <title>Quantitative evaluation of the design of the generator</title> <nav-pointer rid="c3-tab1">84</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 3.2</label> <title>Quantitative comparison between different tileable synthesis methods.</title> <nav-pointer rid="c3-tab2">88</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 4.1</label> <title>Number of trainable parameters in the decoders and amount of latent texture channels for different neural BTF compression algorithms.</title> <nav-pointer rid="c4-tab1">111</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 5.1</label> <title>Results of our ablation study, across a variety of metrics.</title> <nav-pointer rid="c5-tab1">134</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 5.2</label> <title>Model sizes for different methods, and their evaluation time in seconds.</title> <nav-pointer rid="c5-tab2">138</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 5.3</label> <title>Comparisons of our method with previous work on SVBRDF capture</title> <nav-pointer rid="c5-tab3">139</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 5.4</label> <title>Comparisons of our method with previous work for a <italic>curdoroy fabric</italic></title> <nav-pointer rid="c5-tab4">152</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 5.5</label> <title>Comparisons of our method with previous work for a <italic>suede leather</italic></title> <nav-pointer rid="c5-tab5">153</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 6.1</label> <title>Ablation study of the neural architecture and data augmentation.</title> <nav-pointer rid="c6-tab1">166</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 6.2</label> <title>Results for real depth images varying the input.</title> <nav-pointer rid="c6-tab2">168</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 6.3</label> <title>Average correlation with rankings obtained through the tSTE Embedding.</title> <nav-pointer rid="c6-tab3">174</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 7.1</label> <title>Average environment map reconstruction error for different methods.</title> <nav-pointer rid="c7-tab1">203</nav-pointer></toc-entry>
<toc-entry content-type="table"><label>&#x2022; 7.2</label> <title>Average render reconstruction error for different methods.</title> <nav-pointer rid="c7-tab2">206</nav-pointer></toc-entry>
</toc>
</front-matter>
<book-body>
<book-part id="c1" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_xxvi"/><target target-type="page" id="pges_1"/>Chapter 1.</label>
<title>Introduction</title>
</title-group>
</book-part-meta>
<body>
<disp-quote>
<p><italic>&#x201C;We may say most aptly that the Analytical Engine weaves algebraic patterns just as the Jacquard loom weaves flowers and leaves.</italic></p>
<attrib>&#x2014; <bold><italic>Ada Lovelace</italic></bold></attrib>
</disp-quote>
<p>Generating lifelike virtual scenes has been a long-standing goal in computer graphics. This objective is now closer than ever before, thanks to the development of efficient methods for creating such scenes. Realistic virtual scenes have a range of applications, including media and entertainment, such as art, video games, augmented and virtual reality, and films. Furthermore, they are also important in industrial and commercial settings, such as computer-aided design, architecture, and fashion. We illustrate some use cases in <xref ref-type="fig" rid="fig-1.1">Figure 1.1</xref>.</p>
<fig id="fig-1.1">
<label>Figure 1.1:</label>
<caption><title>Renders of virtual scenes in different use cases. (a) A frame of Pixar&#x2019;s <italic>Soul</italic> (2020), directed by Peter Docter. (b) A frame of <italic>The Last of Us Part II</italic> (2020), Naughty Dog. (c) Garments designed with <italic>SEDDI Author</italic>. (d) A render of <italic>Ceramic House</italic>, courtesy of <italic>Studio RAP</italic>.</title></caption>
<alt-text>Renders of virtual scenes in different use cases. (a) A frame of Pixar&#x2019;s Soul (2020), directed by Peter Docter. (b) A frame of The Last of Us Part II (2020), Naughty Dog. (c) Garments designed with SEDDI Author. (d) A render of Ceramic House, courtesy of Studio RAP.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.1.jpg"/>
</fig>
<p>Achieving photo-realism in virtual scenes requires accurate and efficient representations of the optical behavior of materials, object geometry, and illumination. Although many methods have been developed for generating realistic imagery using these representations, producing photo-realistic images often requires a significant amount of manual labor from skilled artists, as well as substantial computational resources. One of the key bottlenecks in this process is the generation of digital representations of materials. Materials, depicted in <xref ref-type="fig" rid="fig-1.2">Figure 1.2</xref>, are ubiquitous in the real world and exhibit an astonishing variety of structures, regularities, and colors. Realistically rendering them also requires modeling how they interact with light, taking into account their levels of glossiness, transparency, roughness, and microgeometric structures, all of <target target-type="page" id="pges_2"/>which impact their appearance. This wide range of material properties poses challenges to the development of efficient representations that can capture all the variations required to render the real world in a visually compelling way.</p>
<fig id="fig-1.2">
<label>Figure 1.2:</label>
<caption><title>Examples of real-world materials and textures.</title></caption>
<alt-text>Examples of real-world materials and textures.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.2.jpg"/>
</fig>
<p>Furthermore, several design and manufacturing processes involve the manual design of materials, the outcome of which remains uncertain until different fabrication steps have been completed. This work methodology is ubiquitous in industries such as garment design or fashion, and it is slow and wastes a significant amount of valuable resources. The fashion industry alone is projected to generate more than 3 million metric tons of CO<sub>2</sub> emissions (10% of the annual worldwide emissions) and waste 170 billion cubic meters of water in 2025 [<xref ref-type="bibr" rid="CIT329">RSS16</xref>], significantly exacerbating the climate emergency. Advanced techniques in engineering and computing have opened up the possibility of digitally creating and editing simple materials, such as plastics or metals. However, complex deformable materials, notably fabrics, have not been extensively studied due to their intricate optical and mechanical properties. This is problematic, as billions of garments are manufactured every year (a staggering 80 billion in 2015 alone [<xref ref-type="bibr" rid="CIT280">Mor15</xref>], a number that has only increased since then). Besides, establishing a relationship between physical parameters that are relevant to the manufacturing process and perceptual parameters that closely align with how designers perceive materials is a challenging task. Achieving this mapping could be instrumental in enabling designers to modify materials according to their creative vision and effectively communicate their intentions to manufacturers.</p>
<p>Although plenty of solutions for material digitization exist, they are often limited in terms of accuracy, scalability, or cost. Accurately generating digital materials can be expensive and challenging, as it requires the use of inaccessible capture devices such as <italic>gonioreflectometers</italic>, or significant manual effort from skilled artists. These resource requirements not only lead to inefficiencies but also restrict the number of individuals who can benefit from material <target target-type="page" id="pges_3"/>digitization. While cost and time constraints are major obstacles to material digitization, data-driven algorithms and low-cost devices like smartphones offer promising solutions.</p>
<p>Offering digital solutions to the challenges of material digitization has the potential to bring significant benefits to various industries and society as a whole. These benefits include more efficient manufacturing processes, reduced waste, a more dynamic economy, and more immersive virtual experiences for art, media, and entertainment. In this thesis, we present innovative methods that address these challenges by leveraging ideas from computer vision, physically-based rendering, and recent advances in machine learning.</p>
<p>In the following sections, we provide a theoretical background of physically-based rendering (<xref ref-type="sec" rid="c1-s1">Section 1.1</xref>), the open problems in the literature (<xref ref-type="sec" rid="c1-s2">Section 1.2</xref>), our goals (<xref ref-type="sec" rid="c1-s3">Section 1.3</xref>) contributions to tackle these problems (<xref ref-type="sec" rid="c1-s3.s1">Section 1.3.1</xref>) and a list of measurable outcomes of the contributions of this thesis (<xref ref-type="sec" rid="c1-s4">Section 1.4</xref>).</p>
<sec id="c1-s1">
<label>1.1</label>
<title>Theoretical Background</title>
<p>In this section, we introduce important theoretical background, which is relevant to the work executed in this thesis. <italic>This background is derived from a survey on deep intrinsic image decomposition [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>], which was written during the development of this thesis.</italic></p>
<sec id="c1-s1.s1">
<label>1.1.1</label>
<title>Physically-Based Rendering</title>
<p>If we look at any simple scene surrounding us, such as the photography in <xref ref-type="fig" rid="fig-1.3">Figure 1.3</xref>, we can find a plethora of optical interactions: indirect lighting (color bleeding), internal scattering in translucent objects, caustics, anisotropic and glossy reflections, etc. Far from the traditional assumptions in intrinsic imaging of diffuse (lambertian) shading, direct lighting, and diffuse albedo materials.</p>
<fig id="fig-1.3">
<label>Figure 1.3:</label>
<caption><title>Example of light transport and material interactions. Secondary bounces of light produce reflection caustics (chrome pen) and color bleeding from the green book. The wax candle exhibits multiple internal (subsurface) scattering of photons. The yellow silk fabric of the book cover shows specular anisotropic reflections due to yarn orientation. Figure from [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>].</title></caption>
<alt-text>Example of light transport and material interactions. Secondary bounces of light produce reflection caustics (chrome pen) and color bleeding from the green book. The wax candle exhibits multiple internal (subsurface) scattering of photons. The yellow silk fabric of the book cover shows specular anisotropic reflections due to yarn orientation. Figure from [Gar+22].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.3.jpg"/>
</fig>
<p>In the following, we provide an overview of the theoretical background behind the image formation model, its derivation for non-diffuse materials, and the link with physically-based and inverse rendering. For a deeper dive into any of the concepts quite briefly described below, we recommend reading the book on physically based rendering by Pharr et al. [<xref ref-type="bibr" rid="CIT307">PJH20</xref>] and the survey by Guarnera et al. [<xref ref-type="bibr" rid="CIT122">Gua+19</xref>].</p>
<p>The color and luminosity at any point of an image, the incoming <italic>irradiance</italic>, is proportional to the sum of the outgoing <italic>radiance</italic> from all the visible points of the scene towards the camera sensor at the corresponding pixel, <italic>I</italic>, resulting <target target-type="page" id="pges_4"/>from multiple interactions between light and matter in the scene. Naturally, this is a simplification: even if we consider the camera lenses and color filters as part of the scene, the interaction of irradiance and the sensor point affects the result, and both electronic and film cameras have specific additional image formation steps which can be simulated. In physically-based rendering, the general approach to compute this value is to use Monte Carlo estimators of the pixel and shading integral [<xref ref-type="bibr" rid="CIT190">Kaj86</xref>], which has the general form of:</p>
<p><disp-formula id="Eq001-1"><label>(1.1)</label> <mml:math id="M6" display='block'><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mo>&#x222B;</mml:mo><mml:mi>&#x03C7;</mml:mi></mml:msub><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mo>)</mml:mo><mml:mi>d</mml:mi><mml:mi>x</mml:mi></mml:math></disp-formula></p>
<p>where <italic>f</italic> is a function that defines the radiance towards pixel <italic>I</italic> and is defined on a domain <inline-formula><mml:math id="M7" display='block'><mml:mi>&#x03C7;</mml:mi></mml:math></inline-formula>, generally a unit sphere, or the set of all the surfaces (<italic>A</italic>) in the scene, and depends on the scene parameters &#x0398;, which might include the definition of the geometry (normals, z-depth, vertices), material (albedo, BRDF), or illumination sources &#x015B; far-field environment lighting, or 3D light emitters (point, area, objects).</p>
<p>The value of <italic>I</italic> is defined for a given <italic>&#x03BB;</italic>, which is the spectral band of the camera sensor. We could define it for as many bands as desired (including non-visible <target target-type="page" id="pges_5"/>ones) but in practice, the majority of camera sensors mimic the human visual system and are commonly three-band: <inline-formula><mml:math id="M8" display='block'><mml:mi>&#x03BB;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo>{</mml:mo><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mi>G</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mo>}</mml:mo></mml:math></inline-formula>. In subsequent rendering equations, we will simplify the notation by assuming a single-band, and omitting the term <italic>&#x03BB;</italic>.</p>
<p><xref ref-type="disp-formula" rid="Eq001-1">Equation 1.1</xref> is an integral of integrals (see <xref ref-type="fig" rid="fig-1.4">Figure 1.4</xref>): to account for all the light arriving at a surface point <italic>p</italic><sub>1</sub> to a sensor pixel at <italic>p</italic><sub>0</sub>, we have to estimate all the contributions of light from all the surfaces of the scene, recursively tracing paths bouncing in surfaces <inline-formula><mml:math id="M9" display='block'><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x2026;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:math></inline-formula> until we reach the light emitted by a source <inline-formula><mml:math id="M10" display='block'><mml:msub><mml:mi>L</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. This is referred as the Light Transport Equation (LTE) in rendering and it is another way of seeing the <xref ref-type="disp-formula" rid="Eq001-1">equation 1.1</xref>. To compute the radiance reaching the pixel <italic>I</italic>, that is from <italic>p</italic><sub>1</sub> to <italic>p</italic><sub>0</sub> in <xref ref-type="fig" rid="fig-1.4">Figure 1.4</xref>, we would need to solve:</p>
<p><disp-formula id="Eq001-2"><label>(1.2)</label> <mml:math id="M11" display='block'><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:munderover><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>with <inline-formula><mml:math id="M12" display='block'><mml:msub><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#xAF;</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> being the radiance scattered over a path <inline-formula><mml:math id="M13" display='block'><mml:msub><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#xAF;</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> with <italic>n</italic> + 1 vertices <inline-formula><mml:math id="M14" display='block'><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x2026;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:math></inline-formula> and computed as:</p>
<p><disp-formula id="Eq001-3"><label>(1.3)</label> <mml:math id="M15" display='block'><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#xAF;</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:munder><mml:mrow><mml:msub><mml:mo>&#x222B;</mml:mo><mml:mi>A</mml:mi></mml:msub><mml:msub><mml:mo>&#x222B;</mml:mo><mml:mi>A</mml:mi></mml:msub><mml:mo>&#x2026;</mml:mo><mml:msub><mml:mo>&#x222B;</mml:mo><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>&#x23DF;</mml:mo></mml:munder><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munder><mml:msub><mml:mi>L</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>T</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#xAF;</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mi>d</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>We can integrate over solid angles in the unit sphere, or surfaces (A) of the scene, being <inline-formula><mml:math id="M16" display='block'><mml:mi>d</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> the differential area at point <italic>p<sub>i</sub></italic>. Note the term <inline-formula><mml:math id="M17" display='block'><mml:mi>T</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#xAF;</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, named <italic>throughput</italic> of the path: the fraction of radiance from the light source that arrives at the camera after all of the scattering at vertices between them. The total transmitted energy will be reduced at each interaction event, as some wavelengths (colors) are absorbed or scattered away from the observer. This is a common trade-off decision in many Monte Carlo rendering engines; longer paths per sample are costly to compute, while contributing less and less energy with each additional vertex, but in some scenes, they might be very relevant to reduce variance (noise) and converge with fewer samples to an accurate image.</p>
<fig id="fig-1.4">
<label>Figure 1.4:</label>
<caption><title>Example of light transport, connecting a light source <italic>p</italic><sub>3</sub> to a pixel <italic>I</italic> at <italic>p</italic><sub>0</sub>. Multiple paths like this one will need to be explored to provide a good statistical estimate of the radiance from <italic>p</italic><sub>1</sub> to <italic>p</italic><sub>0</sub>. Figure from [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>].</title></caption>
<alt-text>Example of light transport, connecting a light source p3 to a pixel I at p0. Multiple paths like this one will need to be explored to provide a good statistical estimate of the radiance from p1 to p0. Figure from [Gar+22].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.4.jpg"/>
</fig>
<p>If there are interactions with non-opaque materials (human skin, cloth, marble) or participating media, such as liquids or smoke, we have to integrate volumetric scattering interactions of photons along the path between the light sources and the camera sensor pixel, requiring a more complex mathematical model such as the Radiative Transport Equation (RTE).</p>
</sec>
<sec id="c1-s1.s2">
<label>1.1.2</label>
<title>Geometry</title>
<p>If we take a look at <xref ref-type="fig" rid="fig-1.4">Figure 1.4</xref>, we can observe that materials are distributed on discrete 3D objects: the table and the cup, and even the light source if <target target-type="page" id="pges_6"/>we consider it as an emissive material. We can assume that one of the most important &#x0398; parameters is geometry, usually in the form of 3D vertices and edges, normals (<italic>N</italic>), or depth maps (<italic>D</italic>). Please note that, in contrast to a a full 3D mesh, a single camera-view depth image (<italic>a.k.a.</italic> Z-buffer) is an incomplete definition because the non-visible surfaces are undefined, and the reflected light paths cannot be traced behind the visible objects. For example, Nimier et al.[<xref ref-type="bibr" rid="CIT297">Nim+19</xref>] require multiple views of a smoke volume in order to reconstruct its 3D density distribution by inverse rendering.</p>
<p>There are additional ways of defining this geometry, such as implicit surfaces, but the material distribution is more complex when there are no clear surface boundaries. That is the case of heterogeneous participating materials: human skin, airlight, mixed liquids, smoke, etc., where light transport between two points has to be in turn evaluated along the path to account for all the possible scattering effects. For instance, imagine a small cloud of vapor between points <italic>p</italic><sub>1</sub> and <italic>p</italic><sub>2</sub> in <xref ref-type="fig" rid="fig-1.4">Figure 1.4</xref>, at each infinitesimal step along the path between those points, there <target target-type="page" id="pges_7"/>will be a possibility of absorption (collision with a water particle) that will reduce the energy, but there is also a possibility of receiving incoming energy (not emitted by <italic>p</italic><sub>2</sub>), as the cloud itself is receiving direct illumination from the light source and multiple scattering events distribute the light across its volume.</p>
<p>In order to compute the rendering equation in volumetric media, a distribution of parameters is required at any point of the space, not only on the 3D surfaces. The usual representations include solid 3D implicit functions (for instance 3D Perlin noise, used in cloud procedural generation), meshes and distance fields, <target target-type="page" id="pges_8"/>or discrete volumetric grids which store the scattering probability function at any point of the space (also known as <italic>phase function</italic>) by means of voxels.</p>
</sec>
<sec id="c1-s1.s3">
<label>1.1.3</label>
<title>Materials</title>
<p>Each time the light interacts with a material, there is a loss of energy and a transformation of the original wavelength reflected towards the observed direction. In rendering, the result depends on the intrinsic material response for those two angles: incident light and viewing direction (<italic>e.g.</italic>: the camera, or another element of the scene). For <italic>surfaces</italic>, this response is modeled with a Bidirectional Reflectance Distribution Function (BRDF) <inline-formula><mml:math id="M18" display='block'><mml:msub><mml:mi>f</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> which yields at each particular 3D surface point x, and for each incident direction <inline-formula><mml:math id="M19" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, the fraction of reflected radiance observed from a direction <inline-formula><mml:math id="M20" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:math></inline-formula>.</p>
<p>The total reflected radiance <italic>L</italic> at any point x can be obtained by integrating with <xref ref-type="disp-formula" rid="Eq001-4">Equation 1.4</xref> over the positive hemisphere &#x03A9;<sup>+</sup>, to sample the whole incident light attenuated by the cosine term (dot product between the incident light direction <italic>L<sub>i</sub></italic> and the normal of the surface) [<xref ref-type="bibr" rid="CIT190">Kaj86</xref>]:</p>
<p><disp-formula id="Eq001-4"><label>(1.4)</label> <mml:math id="M21" display='block'><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mo>&#x222B;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:msub><mml:msub><mml:mi>f</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></disp-formula></p>
<p>Note that by integrating the computed radiance <inline-formula><mml:math id="M22" display='block'><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of the points sampled from <italic>p</italic><sub>0</sub> at the camera sensor (<inline-formula><mml:math id="M23" display='block'><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula>, see <xref ref-type="fig" rid="fig-1.4">Figure 1.4</xref>), we are obtaining the irradiance at the sensor and the corresponding image pixel values (<italic>I</italic> in <xref ref-type="disp-formula" rid="Eq001-1">equation 1.1</xref>). Naturally, even the pixels themselves can be sampled several times and integrated over the camera sensor with another Monte Carlo estimator to minimize aliasing effects.</p>
<p>The BRDF can be extended with a Bidirectional Transmittance Distribution Function (BTDF) to conform a full Bidirectional Scattering Function (BSDF), defined in the full sphere &#x03A9;. The model can be further extended to account for Surface Scattering phenomena (BSSDF).</p>
<p>These functions have a minimum of four dimensions (input-output pair directions in polar coordinates) and usually three RGB values as output. It is thus technically possible to choose a discrete set of (<inline-formula><mml:math id="M24" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:math></inline-formula>) orientations and create a lookup table to interpolate the response of the material, which is captured with multiple light and camera positions (<italic>e.g.</italic>with a gonioreflectometer). Storage becomes a major drawback for tabulated data, which can only be reduced through a significant reduction of quality. Moreover, we are considering only homogeneous surface materials, which is not often the case in actual scenes (<italic>e.g.</italic>, a printed paper). Those spatially-varying values <target target-type="page" id="pges_9"/>(SVBRDF) can be stored in a stack of textures; or a Bidirectional Distribution Texture Function (BTF), increasing the dimensions and size of the table.</p>
<p>Beyond direct compression techniques, the most successful approach in graphics has been the use of analytic N-dimensional functions to approximate the reflectance and scattering distributions. The simplest of them, Lambertian diffuse, and Phong specular shading are also well known in the computer vision community. These functions leverage symmetry (isotropy) to model the material with a few parameters. For instance, a Lambertian material only requires knowing the intrinsic albedo, a sort of base color, while the Phong model requires three additional parameters for the specular component (<italic>e.g.</italic>, shininess).</p>
<p>In the following subsections, we review the most common and simple material assumptions used in recent papers, finalizing with the most sophisticated inverse material models which are starting to be studied in computer graphics.</p>
<sec id="c1-s1.s3.1">
<title>Lambertian Assumption</title>
<p>The Lambertian assumption is the most common material reflectance simplification used to tackle the problem of intrinsic image decomposition. It consists of assuming that the BRDF of a surface is constant in all directions (diffuse) and, consequently, the observed light radiance does not depend on the viewpoint. Therefore, we can omit <inline-formula><mml:math id="M25" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:math></inline-formula> in the surface reflectance model <inline-formula><mml:math id="M26" display='block'><mml:msub><mml:mi>f</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula> used in <xref ref-type="disp-formula" rid="Eq001-4">Equation 1.4</xref>. If the surface is diffuse, then <inline-formula><mml:math id="M27" display='block'><mml:msub><mml:mi>f</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:mfrac></mml:math></inline-formula>, with <inline-formula><mml:math id="M28" display='block'><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:math></inline-formula> denoting the diffuse albedo: the constant ratio of incident light which is reflected in any direction, independently of the viewpoint <inline-formula><mml:math id="M29" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:math></inline-formula>. The image pixel value <italic>I</italic> is then given by:</p>
<p><disp-formula id="Eq001-5"><label>(1.5)</label> <mml:math id="M30" display='block'><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:munder><mml:munder><mml:mfrac><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mi>&#x03C0;</mml:mi></mml:mfrac><mml:mo>&#x23DF;</mml:mo></mml:munder><mml:mi mathvariant="normal">A</mml:mi></mml:munder><mml:munder><mml:munder><mml:mrow><mml:msub><mml:mo>&#x222B;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x23DF;</mml:mo></mml:munder><mml:mi>S</mml:mi></mml:munder></mml:math></disp-formula></p>
<p>The intrinsic model then can be defined as,</p>
<p><disp-formula id="Eq001-6"><label>(1.6)</label> <mml:math id="M31" display='block'><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>S</mml:mi></mml:math></disp-formula></p>
<p>where <italic>S</italic> contains all the shading variations due to the geometry of the local surface w.r.t. light direction. In some cases, the shading image should contain contributions of all the lights in the scene (<italic>S</italic><sub>1</sub> + <italic>S</italic><sub>2</sub> + ... + <italic>S</italic>), which for discrete directional lights, can be deterministically estimated with a linear summation:</p>
<p><disp-formula id="Eq001-7"><label>(1.7)</label> <mml:math id="M32" display='block'><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mn>1</mml:mn><mml:mi>K</mml:mi></mml:munderover><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><target target-type="page" id="pges_10"/>In the case of a more realistic illumination representation, such as environment lighting, or indirect light, the shading computation requires sampling the whole hemisphere &#x03A9;<sup>+</sup>, often recursively sampling other surfaces to approximate the integral of the incoming light. The shading component <italic>S</italic> within <xref ref-type="disp-formula" rid="Eq001-5">Equation 1.5</xref> in the integral form, is difficult to compute and not so easily invertible and differentiable, so until recently, most intrinsic decomposition methods assumed the simpler formula described in <xref ref-type="disp-formula" rid="Eq001-6">Equations 1.6</xref> and <xref ref-type="disp-formula" rid="Eq001-7">1.7</xref>. Note that the illumination visibility is not considered in most cases (E.g.: cast shadows).</p>
</sec>
<sec id="c1-s1.s3.2">
<title>Non-Lambertian Assumption</title>
<p>There are two possible sources producing a Lambertian shading, either a surface which has an extremely rough micro-geometry, and thus reflects light equally in multiple random directions at any differential patch of the surface, or a very diffuse light source, coming from any direction with equal intensity (<italic>e.g.</italic>, a foggy day). Both scenarios can be combined (<xref ref-type="disp-formula" rid="Eq001-4">Equation 1.4</xref>): the shiniest object on a foggy day will look quite diffuse, while even the most diffuse materials tend to project specular reflections under focused lighting from certain view angles. However, the majority of materials in the world are not Lambertian: even the most diffuse surface will exhibit Fresnel reflections when observed at grazing angles. Therefore, most surfaces will show the view-dependent effects that are classified as specular reflections. This separation between specular and Lambertian is rather pragmatic, but arbitrary, as even a simple microfacet model (shown in <xref ref-type="fig" rid="fig-1.5">Figure 1.5</xref>) requires multiple analytic 3D lobes to approximate the 3D reflectance response for an infinitesimal incoming light ray <inline-formula><mml:math id="M33" display='block'><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:math></inline-formula>). The term <italic>specular</italic> is usually applied to narrow lobes with high probability of scattering radiance, producing high luminance values at pixels (highlights). This family of materials are of the general form:</p>
<p><disp-formula id="Eq001-8"><label>(1.8)</label> <mml:math id="M34" display='block'><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mo>&#x222B;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:msub><mml:munder><mml:munder><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>&#x23DF;</mml:mo></mml:munder><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:munder><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></disp-formula></p>
<p><disp-formula id="Eq001-9"><label>(1.9)</label> <mml:math id="M35" display='block'><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>where <italic>f<sub>NL</sub></italic> is a non-lambertian BRDF composed by two components: <italic>f<sub>d</sub></italic>, a diffuse isotropic lobe, and <italic>f<sub>s</sub></italic>, a specular lobe which depends on the camera viewpoint <inline-formula><mml:math id="M36" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:math></inline-formula>.</p>
<fig id="fig-1.5">
<label>Figure 1.5:</label>
<caption><title>2D depiction of a physically-principled BSDF theorical model. In most standard representations, the continuous reflectance 4D function is discretized into a combination of analytic lobes (Cosine, GGX) which can be easily computed sampled. Note that subsurface scattering (photons traveling through the medium) is not composed by multiple lobes. They are depicted for descriptive purposes, but it is rather approximated by a constant value or a diffusion profile, if not explicitly computed by simulating multiple scattering events with path tracing or photon mapping. Figure from [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>].</title></caption>
<alt-text>2D depiction of a physically-principled BSDF theorical model. In most standard representations, the continuous reflectance 4D function is discretized into a combination of analytic lobes (Cosine, GGX) which can be easily computed sampled. Note that subsurface scattering (photons traveling through the medium) is not composed by multiple lobes. They are depicted for descriptive purposes, but it is rather approximated by a constant value or a diffusion profile, if not explicitly computed by simulating multiple scattering events with path tracing or photon mapping. Figure from [Gar+22].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.5.jpg"/>
</fig>
<p><bold>Dichromatic Reflection Model</bold> This particular Non-Lambertian model [<xref ref-type="bibr" rid="CIT264">MFS08</xref>; <xref ref-type="bibr" rid="CIT389">Tom94</xref>] separates the object in two reflection components (<italic>S</italic><sub>d</sub>, <italic>S</italic><sub>s</sub>), but considers that the specular component <italic>S</italic><sub>s</sub> might have a color &#x03B1;<sub>s</sub> which could differ from the color of the reflected light:</p>
<p><disp-formula id="Eq001-10"><label>(1.10)</label> <mml:math id="M37" display='block'><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:munder><mml:munder><mml:mrow><mml:msub><mml:mo>&#x222B;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:msub><mml:msub><mml:mi>f</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x23DF;</mml:mo></mml:munder><mml:msub><mml:mi>S</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:munder></mml:math></disp-formula></p>
<p><disp-formula id="Eq001-11"><label><target target-type="page" id="pges_11"/>(1.11)</label> <mml:math id="M38" display='block'><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:munder><mml:munder><mml:mrow><mml:msub><mml:mo>&#x222B;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:msub><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x23DF;</mml:mo></mml:munder><mml:msub><mml:mi>S</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:munder></mml:math></disp-formula></p>
<p><disp-formula id="Eq001-12"><label>(1.12)</label> <mml:math id="M39" display='block'><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:msub><mml:mi>S</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:msub><mml:mi>S</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></disp-formula></p>
<p>This is the case of metallic materials which, unlike dielectric ones, will show specular reflections with a change in wavelength. Some additional effects such as colored interreflections might be also captured in all layers.</p>
<p><bold>Phong and Blinn-Phong.</bold> The dichromatic model can be extended with one of the most adopted analytic approximations, either Phong (<inline-formula><mml:math id="M40" display='block'><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>L</mml:mi></mml:mrow><mml:mi>P</mml:mi></mml:msubsup></mml:math></inline-formula>) or Blinn-Phong (<inline-formula><mml:math id="M41" display='block'><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>L</mml:mi></mml:mrow><mml:mi>P</mml:mi></mml:msubsup></mml:math></inline-formula>), which could be estimated with Monte Carlo integration and arbitrary lighting, or analytically computed with directional light sources:</p>
<p><disp-formula id="Eq001-13"><label>(1.13)</label> <mml:math id="M42" display='block'><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mi>L</mml:mi></mml:mrow><mml:mi>P</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>k</mml:mi></mml:msup></mml:math></disp-formula></p>
<p><disp-formula id="Eq001-14"><label>(1.14)</label> <mml:math id="M43" display='block'><mml:msubsup><mml:mi mathvariant="normal">f</mml:mi><mml:mi>NL</mml:mi><mml:mi>BP</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B1;</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi mathvariant="normal">N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03B1;</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="normal">r</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi mathvariant="normal">v</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi mathvariant="normal">k</mml:mi></mml:msup></mml:math></disp-formula></p>
<p>where &#x03B1;<sub>d</sub> and &#x03B1;<sub>s</sub> are the colors of the diffuse and the specular reflections, the halfway vector <inline-formula><mml:math id="M44" display='block'><mml:mi mathvariant="normal">h</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:mfrac></mml:math></inline-formula> depends on the light direction <italic>w</italic><sub>i</sub> and the view direction <italic>w<sub>o</sub></italic>. The size of the specular lobe is determined by the scalar term <italic>k</italic> &#x2208; &#x211D;.</p>
</sec>
<sec id="c1-s1.s3.3">
<title>Beyond Dichromatic Models: Physically-based Materials</title>
<p>Naturally, the breadth of materials that can be synthesized with the previous models is very limited and not quite realistic in most cases. The advent of physically-based materials has introduced many variations [<xref ref-type="bibr" rid="CIT154">Hil+15</xref>] of the original microfacets models [<xref ref-type="bibr" rid="CIT391">TS67</xref>], which assume that a surface is composed of many very tiny facets that reflect light perfectly. By controlling the statistical distribution of their orientations, the roughness of the surface varies from mirror-like to almost diffuse. Additional optical properties are introduced in these models: Fresnel view-dependent reflectivity, multiple specular lobes, metalness (conductive materials such as gold, change the color of the highlights), multiple reflection and refraction lobes, etc.</p>
<p>The separation of the lobes described in the multi-lobed physically-based model shown in <xref ref-type="fig" rid="fig-1.5">Figure 1.5</xref> is not arbitrary, grouping reflections and refractions which share the same orientation and energy level. For instance, the main diffuse lobe is grouping multiple different orientations which are not view-dependent and share the same intensity and color. If it covers the full hemisphere with a cosine-like ratio, it is often referred as <italic>Lambertian</italic>. Likewise, the specular transmitted lobe is grouping a view-dependent peak that the observer would only see when the translucent surface is between the light source and the camera. <target target-type="page" id="pges_12"/>Even if it is often called a <italic>single scatter</italic> lobe, likely multiple internal scattering bounces of light are also included in this group.</p>
</sec>
</sec>
<sec id="c1-s1.s4">
<label>1.1.4</label>
<title>Illumination</title>
<p>The illumination is a significant contributor to the shading term (<italic>S</italic>) in most decompositions. From a rendering perspective, as shown in <xref ref-type="disp-formula" rid="Eq001-15">Equation 1.15</xref>, the computation of the pixel radiance, <italic>L</italic>, requires considering both the emitters <italic>L</italic><sub>e</sub> and the irradiance: the integral of all the incoming lighting at the observed point. This incoming illumination is often neglected, considering only point light or directional analytic emitters, which simplify the shading computation by removing <italic>L</italic><sub>i</sub> from the integral. If <italic>f</italic><sub>r</sub> is also assumed to be Lambertian, only the form factor given by the cosine of the surface normal, and the light direction remains.</p>
<p><disp-formula id="Eq001-15"><label>(1.15)</label> <mml:math id="M45" display='block'><mml:mtable columnspacing="0em" columnalign="right left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mo>+</mml:mo><mml:msub><mml:mo>&#x222B;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>+</mml:mo></mml:msup></mml:msub><mml:msub><mml:mi>f</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>N</mml:mi><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In actual scenes, the incoming lighting is a combination of emitted or reflected illumination from distant objects (far field) and local surfaces close to the observed area (near field). The former is usually approximated in computer graphics with environment lighting, often based in High-Dynamic-Range (HDR) <target target-type="page" id="pges_13"/>images mapped into an infinite sphere or cube surrounding the scene, while the latter can be derived from the far field illumination, by simulating the local secondary light bounces. If the geometry does not change, an environment map can be stored at multiple scene locations and distances, to include near field effects more accurately (Spatially Varying Environment Maps), although at a great memory cost, and only for static scenes.</p>
<p>To reduce the sampling and size of environment maps and simplify the computation, Ramamoorthi and Hanrahan [<xref ref-type="bibr" rid="CIT319">RH01a</xref>] proposed their compression with Spherical Harmonics (SH), a set of orthonormal basis functions defined on the spherical domain (elevation &#x03B8; and azimuth &#x03D5; angles). Thus the <xref ref-type="disp-formula" rid="Eq001-16">equation 1.16</xref> describes the irradiance <italic>E</italic> as the sum of bases weighted by the cosine decay term <italic>A</italic><sub>&#x03B8;</sub> and the illumination coefficient <italic>L</italic><sub>&#x03B8;,</sub>. By changing the representation of the bases <inline-formula><mml:math id="M46" display='block'><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> to polynomial coordinates of a unit normal <inline-formula><mml:math id="M47" display='block'><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>z</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:math></inline-formula>, this becomes an efficient vector dot product operation (<xref ref-type="disp-formula" rid="Eq001-17">Equation 1.17</xref>). With the required modifications, this strategy is feasible with other orthogonal basis functions on the sphere.</p>
<p><disp-formula id="Eq001-16"><label>(1.16)</label> <mml:math id="M48" display='block'><mml:mi>E</mml:mi><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mover><mml:mi>A</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>l</mml:mi></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="Eq001-17"><label>(1.17)</label> <mml:math id="M49" display='block'><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mover><mml:mi>Y</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>T</mml:mi></mml:msup><mml:mi>L</mml:mi></mml:math></disp-formula></p>
<p><disp-formula id="Eq001-18"><label>(1.18)</label> <mml:math id="M50" display='block'><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>L</mml:mi></mml:math></disp-formula></p>
<p>If we want to account for near field occlusion and interreflection, it is possible to precompute those local interactions, because they depend on the object geometry and materials, and not on far field illumination. This family of techniques is known as <italic>precomputed radiance transfer</italic> (PRT) [<xref ref-type="bibr" rid="CIT368">SKS02</xref>]: they precompute multiple events of light transport (see <xref ref-type="fig" rid="fig-1.4">Figure 1.4</xref>) into the <italic>T</italic> term in <xref ref-type="disp-formula" rid="Eq001-18">Equation 1.18</xref> with Monte Carlo pathtracing. In this fashion, each pixel will have secondary light bounces stored in a light transport map. If only the visibility term <italic>V</italic> (<italic>w</italic><sub>i</sub>) is considered for, the method will be storing the ambient occlusion shadows, but not colored interreflections.</p>
<p>In <xref ref-type="fig" rid="fig-1.6">Figure 1.6</xref>, we can see a pyramid of spherical harmonics bases <italic>Y</italic><sub>&#x03B8;,</sub> with different coefficients. It is important to know that, although usually nine bases are considered enough to account for 99% of the far-field irradiance at diffuse surfaces, this percentage is significantly smaller in glossy surfaces (requiring many more coefficients). Moreover, a small number of coefficients will never account for high frequency effects, such as cast shadows from high-frequency light sources (E.g: a point light representing the sun), even producing ringing artifacts if we try to increase the accuracy by adding more bases. There are other popular basis in rendering such as Haar Wavelets or Spherical Gaussians (SG) [<xref ref-type="bibr" rid="CIT411">Wan+09</xref>], which also have very interesting properties.</p>
<fig id="fig-1.6">
<label>Figure 1.6:</label>
<caption><title>Example of spatially varying illumination encoding with Spherical Harmonics (SH). The incoming lighting can be computed globally for the whole scene (far field), or locally, at multiple points (near field) as we show for the two samples near each colored wall. If we project the irradiance (top row) into an SH basis we obtain a diffuse low-frequency representation (examples in bottom row). Figure from [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>].</title></caption>
<alt-text>Example of spatially varying illumination encoding with Spherical Harmonics (SH). The incoming lighting can be computed globally for the whole scene (far field), or locally, at multiple points (near field) as we show for the two samples near each colored wall. If we project the irradiance (top row) into an SH basis we obtain a diffuse low-frequency representation (examples in bottom row). Figure from [Gar+22].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.6.jpg"/>
</fig>
</sec>
</sec>
<sec id="c1-s2">
<label><target target-type="page" id="pges_14"/>1.2</label>
<title>Open Problems</title>
<p>In this section, we aim to address the challenges inherent in creating precise and efficient virtual representations of materials and scene radiance. We begin by discussing the current difficulties in acquiring and editing high-quality digital materials that are suitable for photo-realistic physically based rendering. Specifically, we will explore the challenges associated with tileable texture synthesis and image-based material propagation.</p>
<p>We then turn our attention to the major roadblocks that hinder low-cost digitization of the mechanical and optical parameters of real-world materials. We will examine the various obstacles to overcome when attempting to accurately capture and reproduce the physical characteristics of materials in a virtual environment, using scalable and affordable pipelines.</p>
<p>Lastly, we will review the limitations of current research on radiance encoding using neural networks, with a particular focus on representing material reflectance and scene illumination. By identifying these challenges, we hope to provide insight into the current state-of-the-art in material and radiance modeling for physically based rendering.</p>
<sec id="c1-s2.1">
<title>Digital Material Generation</title>
<p>Accurate and high resolution material reflectance digitization is a crucial problem for generating compelling and realistic virtual environments. As we have mentioned in <xref ref-type="sec" rid="c1-s1.s3">Section 1.1.3</xref>, this digitization process is typically done using a <italic>gonioreflectometer</italic>, a highly complex machine designed to measure the photometric response of a material using light sources and cameras placed at different positions around the hemisphere. Existing devices for measuring material appearance in spatially-varying samples are limited to a single scale, either micro or mesoscopic. This is a practical limitation when the material has a complex multi-scale structure, like is common on many real-world materials. In many materials, accurately measuring their spatially-varying reflectance requires high resolution microscopic images. For example, the optical behavior of textiles is very dependent on small fibers or their yarn twist [<xref ref-type="bibr" rid="CIT040">CLA19</xref>]. Moreover, many materials also show large-scale <italic>mesoscopic</italic> variations, like prints, plaid, or tartan patterns, which cannot be captured using microscopic photographs. Further, many of such devices cannot capture important optical properties, like the material transmittance, which is key for the realistic rendering of objects like foliage or fabrics. There are no available capture devices or algorithms for accurate and high resolution SVBSDF digitization of heterogeneous materials. In order to introduce such materials, which are very common in many industries like fashion, textile, or leather manufacturing, into rendering pipelines, these machines need to be created and accurate algorithms need to be developed to handle the data <target target-type="page" id="pges_15"/>they capture. Finally, such a device may serve as a ground truth data generation source of training models that operate on more limited data scenarios, like in single image digitization</p>
<fig id="fig-1.7">
<label>Figure 1.7:</label>
<caption><title>Schemes and photos of SEDDI&#x2019;s optical device, which can capture highly accurate and detailed SVBSDFs. (a) Schema of a cross-section of the hemisphere including micro camera and one polar camera. (b) (Top) Schema of the collimated LED design to account for the polarizer and collimating lens; (Bottom) Schema of the main holder and backlight support. (c) Microscopic optical setup. (d) (Top) A Polar Camera; (Bottom) Mid-distance Camera. (e) Interior of the dome with the diffuse LED strip activated. (f) Exterior of the dome and wiring. (g) Holder and exterior cover. Figure from [<xref ref-type="bibr" rid="CIT103">Gar+23</xref>].</title></caption>
<alt-text>Schemes and photos of SEDDI&#x2019;s optical device, which can capture highly accurate and detailed SVBSDFs. (a) Schema of a cross-section of the hemisphere including micro camera and one polar camera. (b) (Top) Schema of the collimated LED design to account for the polarizer and collimating lens; (Bottom) Schema of the main holder and backlight support. (c) Microscopic optical setup. (d) (Top) A Polar Camera; (Bottom) Mid-distance Camera. (e) Interior of the dome with the diffuse LED strip activated. (f) Exterior of the dome and wiring. (g) Holder and exterior cover. Figure from [Gar+23].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.7.jpg"/>
</fig>
<p><bold>Tileable SVBRDFs</bold> Even if reflectance is accurately measured for a large portion of a material, in rendering settings, it is also important that the digital material is <italic>tileable</italic>. A tileable texture, also known as a seamless texture, is a texture image that can be repeated infinitely in all directions without any visible seams or discontinuities. In other words, a tileable texture is designed in a way that it seamlessly connects with identical copies of itself, which can be arranged side-by-side or stacked on top of one another to create a larger, continuous pattern without any visible repetition. This is often used in computer graphics, especially for creating textures on 3D models, or even for tiling backgrounds on websites or applications. We illustrate the importance of <italic>tileable SVBRDFs</italic> in <xref ref-type="fig" rid="fig-1.8">Figure 1.8</xref>. Recent progress has been made in generating tileable textures from a single input example [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>; <xref ref-type="bibr" rid="CIT341">Rod+19</xref>; <xref ref-type="bibr" rid="CIT279">Mor+20</xref>], however, these works present several shortcomings for tileable SVBRDF generation. First, they either assume a particular level of regularity, or the generated textures lose a significant amount of visual fidelity with respect to the input exemplars. Further, these methods are limited to generating a single texture map (i.e. the albedo of the material). This is problematic for tileable SVBRDF synthesis, which requires transforming every texture map in the SVBRDF into a tileable map, all while preserving pixel-wise coherence between maps. All in all, generating high-quality tileable SVBRDFs from a single input example is not currently possible with the methods proposed in previous work.</p>
<fig id="fig-1.8">
<label>Figure 1.8:</label>
<caption><title>On the left, an input, unprocessed SVBRDF, which is not tileable. On the right, a tileable version of the same material. Renders generated using SEDDI&#x2019;s <uri xlink:href="http://www.Textura.ai">Textura.ai</uri>.</title></caption>
<alt-text>On the left, an input, unprocessed SVBRDF, which is not tileable. On the right, a tileable version of the same material. Renders generated using SEDDI&#x2019;s Textura.ai.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.8.jpg"/>
</fig>
<p><bold>Material Attribute Transfer</bold> Effective material representations require understanding the properties that uniquely define them, which we refer to as <italic>visual <target target-type="page" id="pges_16"/>material attributes</italic>. These attributes are spatially-varying parameters that maintain spatial coherency with respect to the material structure, while remaining invariant to changes in the scene illumination or the geometry of the underlying object. For example, they may represent optical properties of a SVBRDF, artistic stylizations or higher-level properties, as in semantic segmentation masks. Obtaining these attributes for large samples of the material is very challenging. One way to address this problem is to obtain them at a small exemplar of the material, and <italic>transferring</italic> them into larger portions of it, using a large input image which serves as <italic>guidance</italic>. Different visual attribute transfer methods have been proposed in the past, however, they are limited in their efficiency, robustness, generalization or overall performance. Ideally, such a transfer method should be robust to new material illumination conditions, camera degradations or input distortions; as well as being efficient (providing interactive transfer times even in low computational budgets), controllable and predictable, and capable of generating high resolution outputs, which are required for realistic SVBRDFs. Developing a method with these characteristics may prove useful for many downstream tasks, like efficient data generation or material digitization pipelines.</p>
</sec>
<sec id="c1-s2.2">
<title>Scalable Digitization</title>
<p><bold>Material Reflectance Estimation</bold> As mentioned, accurate material digitization typically requires expensive devices which take significant amounts of <target target-type="page" id="pges_17"/>time and manual input to operate. One such case is the capture machine illustrated in <xref ref-type="fig" rid="fig-1.7">Figure 1.7</xref>, which is presented in [<xref ref-type="bibr" rid="CIT103">Gar+23</xref>]. These provide highly realistic digital representations of material reflectance at the cost of scalability, hindering their applicability in many real world settings. For instance, many manufacturing processes, like those in the textile industry, generate a massive variety of different materials, which cannot be easily nor cheaply digitized with the required cadence that characterizes these industries. Further, the limited availability of these devices introduces additional inefficiencies and economic and ecological costs, as physical samples of the materials must be shipped from the client to the institution holding the capture device.</p>
<p>Data-driven material reflectance estimation from low-cost devices is a long-standing problem in the literature [<xref ref-type="bibr" rid="CIT004">AWL13</xref>; <xref ref-type="bibr" rid="CIT005">AWL15</xref>; <xref ref-type="bibr" rid="CIT065">Des+18</xref>; <xref ref-type="bibr" rid="CIT148">Hen+21</xref>; <xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT400">VPS21</xref>]. These methods promise to provide scalable, efficient, inexpensive on-site material digitization. They typically rely on one or more flash-lit images of a material taken with a smartphone camera, and a deep learning model trained on synthetic materials, which outputs plausible estimations. However, these approaches present several drawbacks that make them unsuitable for practical digitization workflows: First, many of such methods rely on generative models which tend to produce artifacts, or use large neural networks which puts limits on the output resolution, hindering the quality of the estimated materials. Besides, smartphone flash-lit images introduce significant calibration challenges; while most existing datasets used to train these models are purely synthetic, further limiting the quality of these estimations. Moreover, many of these methods use perceptual losses for training or in test-time optimization, which, while they help digitize stochastic materials, present additional challenges for guaranteeing the repeatability and consistency required for building a digital inventory. On top of these limitations, few-shot material reflectance estimation is still an ill-posed problem, as such, many different outputs may be plausible given the same input. So far, there is no way to quantify uncertainty for this problem, which hinders the applicability of these systems in real-world scenarios, and limits their trustworthiness, efficiency, and reliability.</p>
<p><bold>Fabric Mechanical Behaviour</bold> Besides material reflectance, another important component of scenes is object geometry. While many real-world objects have static shapes, which can be scanned with relatively accessible devices like the LiDAR sensor in <italic>iPhones</italic>, there are materials that change shape when external forces are applied to them, like gravity. In this sense, <italic>deformable materials</italic> require particular computational models and capture pipelines. One of the most common of such materials is fabrics, which are ubiquitous in the real world, and used in our clothing, furniture, or vehicles. For them, accurately capturing their static shape is not enough: We need to be able to simulate their behavior on new scenarios, like new garments or dynamic forces, like wind [<xref ref-type="bibr" rid="CIT029">Ber+23</xref>] or human movement [<xref ref-type="bibr" rid="CIT352">SOC22</xref>]. Furthermore, textiles are incredibly varied in their fabrication patterns (eg different weave patterns or knitted structures), compositions, <target target-type="page" id="pges_18"/>and finishing, which not only determine their optical appearance [<xref ref-type="bibr" rid="CIT040">CLA19</xref>] but also their mechanical behavior, as shown in <xref ref-type="fig" rid="fig-1.9">Figure 1.9</xref>. The problem of capturing fabric mechanical properties so they can be re-simulated in virtual settings has been amply studied in the literature, however, current solutions require tedious human intervention or expensive capture devices, hindering their scalability. Further, there is a lack of understanding of the perceptual similarity of fabric mechanical behavior, which creates additional challenges for designing accurate fabric mechanical parameter estimation pipelines.</p>
<fig id="fig-1.9">
<label>Figure 1.9:</label>
<caption><title>Two real garments and their <italic>digital twins</italic>, for two fabrics with different mechanical behavior, as seen on their final drapes. Figure from [<xref ref-type="bibr" rid="CIT340">Rod+23b</xref>].</title></caption>
<alt-text>Two real garments and their digital twins, for two fabrics with different mechanical behavior, as seen on their final drapes. Figure from [Rod+23b].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.9.jpg"/>
</fig>
</sec>
<sec id="c1-s2.3">
<title>Neural Radiance Encoding</title>
<p><bold>Reflectance</bold> In <xref ref-type="sec" rid="c1-s1.s3">Section 1.1.3</xref>, we have mentioned several limitations of current representations of material reflectance. In particular, Bidirectional Texture Functions (BTFs) provide dense and accurate material reflectance measurements, however, they are typically prohibitive in terms of computational and memory cost. Recent work [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>; <xref ref-type="bibr" rid="CIT316">Rai+20</xref>; <xref ref-type="bibr" rid="CIT220">Kuz+21</xref>] introduce different neural compression algorithms for BTFs, achieving remarkable realism and efficiency. Neural reflectance encoding methods trained on either synthetic or measured BTFs have thus shown promising results for increased realism of rendered materials. However, existing neural material encodings are immutable, meaning that their output for a certain query of UVs, camera, and light vector is fixed once they are trained. This can become very limiting when the fragment of the material used for training is too small or not tileable, which frequently happens when the material has been scanned with a capture device. Therefore, current neural reflectance encodings, while accurate and efficient, have <target target-type="page" id="pges_19"/>severe limitations in terms of editing capabilities, which hinder their applicability and usefulness in real world rendering settings.</p>
<p><bold>Illumination</bold> Finally, as we mention in <xref ref-type="sec" rid="c1-s1.s4">Section 1.1.4</xref>, there are many available approximations for representing real-world illumination in virtual scenes, including Spherical Harmonics, Spherical Gaussians, or Haar Wavelets. These, however, struggle to efficiently and accurately represent non-diffuse illumination conditions, like high-frequency light sources, as the sun or light bulbs. Recent work on neural representations of natural illumination [<xref ref-type="bibr" rid="CIT105">GES22</xref>] provides better approximations for such cases, but struggle with indoor lighting or night scenes with multiple light sources. Furthermore, these neural approximations do not provide sampling and PDF evaluation capabilities, which are essential for Monte Carlo rendering in path tracing. To the best of our knowledge, there are no lighting approximation methods that provide these sampling capabilities, while working accurately on any type of input environment map. Recent work on invertible generative models and implicit neural representations may provide a promising pathway to achieve these objectives.</p>
</sec>
</sec>
<sec id="c1-s3">
<label>1.3</label>
<title>Goals &#x0026; Contributions</title>
<p>The primary objective of this thesis is to create innovative algorithms for computational representation and capture of materials and scene radiance. These are important for several reasons. These digital representations may be used in a wide variety of applications and problems, including computer graphics, video games, film production, virtual and augmented reality, or computer-aided design. They have the potential to significantly increase the visual realism and quality of the virtual environments created in these settings. Furthermore, the accuracy and efficiency of these methods directly impact the quality and computational resources required for these tasks. As such, obtaining accurate and efficient <target target-type="page" id="pges_20"/>representations and digitization algorithms for materials and other scene components can help reduce the computational cost of rendering, design, capture, and modeling, which can lead to faster, more efficient, and scalable pipelines. Moreover, they may enable efficient procedural dataset creation for training models to solve downstream tasks, like dense segmentation, image or video generation, or embodied vision. Finally, besides efficiency and accuracy, additional features like differentiability or uncertainty quantification capabilities create opportunities for inverse rendering applications or active learning.</p>
<p>To this end, we introduce new learning-based methods to address the open problems described before. Our proposed solutions strive to achieve maximum computational and data efficiency, as well as accuracy, control and reliability. Furthermore, our algorithms are designed to be beneficial to end-users by incorporating human requirements such as perceptual components, edition, predictability, robustness, and uncertainty quantification. All of our proposed methods leverage neural networks, incorporating recent advances in implicit representations, architecture and loss function design, and training procedures. They are all fully differentiable, and may be incorporated into optimization pipelines to solve inverse rendering problems. Our solutions for radiance encoding may be seamlessly incorporated into path-tracer engines for increased render efficiency, or into neural scene representations for increased realism. Further, our high quality casual digitization methods allow for scalable and low cost material capture, which can help end users create their own realistic virtual materials, or create large datasets at a lower cost. We provide an overview of this thesis in <xref ref-type="fig" rid="fig-1.10">Figure 1.10</xref>.</p>
<fig id="fig-1.10">
<label>Figure 1.10:</label>
<caption><title>Overview of the structure of this thesis. We propose different algorithms to solve problems (each column) in different components (each color) of virtual scenes.</title></caption>
<alt-text>Overview of the structure of this thesis. We propose different algorithms to solve problems (each column) in different components (each color) of virtual scenes.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.10.jpg"/>
</fig>
<p>Our goals in this thesis tackle the aforementioned challenges in digital material generation, scalable digitization, and radiance encoding, and can be summarized as follows:</p>
<list list-type="bullet">
<list-item><p>Design new learning-based methods for predictable, controllable and high quality digital material propagation, generation, edition, and synthesis.</p></list-item>
<list-item><p>Introduce novel efficient and differentiable representations for scene radiance.</p></list-item>
<list-item><p>Democratize high-quality material capture by presenting new methods which are low cost, scalable, accurate, reliable, and perceptually-validated.</p></list-item>
</list>
<sec id="c1-s3.s1">
<label>1.3.1</label>
<title>Contributions</title>
<p>The work developed in this thesis has led to the following contributions to achieve the goals presented above:</p>
<list list-type="bullet">
<list-item><p><target target-type="page" id="pges_21"/>A deep-learning based method for propagating spatially-varying attributes of a material to larger samples of the same or similar materials. At the core of this method is a lightweight fully-convolutional neural network trained using a <italic>photometric dataset</italic> and an extensive data augmentation policy. These contributions allow our models to generalize to new illumination, color, and geometric conditions. We show the effectiveness of our system on transferring attributes of different types, including surface normals, semantic segmentation, and artistic editions (<xref ref-type="book-part" rid="c2">Chapter 2</xref>).</p></list-item>
<list-item><p>An extension to the previous method, for enabling a dual-scale capture system capable of digitizing a single material at high-resolution and accuracy levels. To do so, we extend the method presented in <xref ref-type="book-part" rid="c2">Chapter 2</xref> for allowing to transfer multiple property maps at the same time, and with a higher degree of accuracy (<xref ref-type="sec" rid="c2-sC">Chapter 2.C</xref>).</p></list-item>
<list-item><p>A generative model capable of creating <italic>tileable textures</italic> using a single image as input. By leveraging state-of-the-art generative adversarial networks (GANs) for texture synthesis, a novel latent space manipulation algorithm, and using the discriminator as a quality metric, our proposed model can generate tileable textures of higher perceptual quality, and at a lower cost than previous work. Further, we propose an extension for generating tileable SVBRDFs from a single input, and show the effectiveness of our method across textures of different levels of regularity. (<xref ref-type="book-part" rid="c3">Chapter 3</xref>).</p></list-item>
<list-item><p>A neural field representation capable of <italic>BTF encoding and transfer</italic>. Building upon previous work on neural material representations, we propose a lightweight latent representation that can be decoded by an implicit neural network into per-texel reflectance values. This latent representation is estimated by an autoencoder-like network, which can be used to propagate BTF values to novel structures, allowing for neural material edition and transfer, and providing new capabilities for these types of representations. (<xref ref-type="book-part" rid="c4">Chapter 4</xref>).</p></list-item>
<list-item><p>A novel single-image capture system capable of generating high-resolution SVBRDFs of materials captured from commodity scanners. Leveraging a custom-built dataset and a novel attention-enhanced image-to-image translation generative model trained with a variety of loss functions, our model provides artifact-free, highly accurate digitizations, using solely microgeometry patterns as cues. We further propose the first uncertainty quantification algorithm for SVBRDF estimation methods, building upon Bayesian deep learning approximations using Monte Carlo dropout during test time, and perceptual BRDF metrics. We show the effectiveness of our uncertainty metric for predicting digitization error at test time and for an active learning experiment which shows that uncertainty sampling helps build more data-efficient capture systems. (<xref ref-type="book-part" rid="c5">Chapter 5</xref>). We illustrate the capabilities introduced by this model in <xref ref-type="fig" rid="fig-1.11">Figure 1.11</xref>.</p></list-item>
<list-item><p><target target-type="page" id="pges_22"/>A learning-based system for digitizing mechanical properties of fabrics using a casual capture setup and a commodity depth camera. Training solely on synthetic data and leveraging an attention-enhanced multi-image neural network, transfer learning, and an extensive data augmentation policy, our proposed method can accurately estimate mechanical properties of fabric samples, without the need for expensive capture equipment. Further, we propose a novel image-based perceptual metric to measure mechanical similarity between fabrics, which we validate with a user study. (<xref ref-type="book-part" rid="c6">Chapter 6</xref>).</p></list-item>
<list-item><p>A novel neural method for environment maps, which provides an efficient global illumination representation. At the core of our method lie two neural networks: a <italic>normalizing flow</italic>, capable of learning the <italic>pdf</italic> of an input environment map, and to efficiently sample directions from it; and a <italic>implicit neural network</italic> with sinusoidal activations capable of mapping from directions to linear RGB radiance values. We show the effectiveness of our method on <italic>Multiple Importance Sampling</italic> applications in rendering, obtaining accurate illumination at a significantly lower cost than traditional methods; as well as better quality environment maps than previous work on illumination representations. (<xref ref-type="book-part" rid="c7">Chapter 7</xref>).</p></list-item>
<list-item><p>A comprehensive study of the capabilities and limitations of deep neural networks for <italic>intrinsic image decomposition</italic>. With a careful review of previous work, approaches, assumptions, datasets, losses, and neural network architectures, we provide a new categorization of these works. We further propose future research directions, building upon recent work on differentiable and inverse rendering and advances In machine learning systems. This contribution has been published in [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>], a part of it is used in <xref ref-type="sec" rid="c1-s1">Section 1.1</xref>.</p></list-item>
</list>
<fig id="fig-1.11">
<label>Figure 1.11:</label>
<caption><title>An illustration of the material digitization process enabled by some methods presented in this thesis. We show a picture of a user casually scanning a piece of fabric (left), the estimated SVBRDF for that material (middle), and a digital render of this SVBRDF (right). Assets generated using SEDDI <uri xlink:href="http://www.Textura.ai">Textura.ai</uri>.</title></caption>
<alt-text>An illustration of the material digitization process enabled by some methods presented in this thesis. We show a picture of a user casually scanning a piece of fabric (left), the estimated SVBRDF for that material (middle), and a digital render of this SVBRDF (right). Assets generated using SEDDI Textura.ai.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.11.jpg"/>
</fig>
<p><target target-type="page" id="pges_23"/>Each of these contributions has been developed in collaboration with other researchers and engineers, many of whom are listed as authors of the publications listed below. However, most of this thesis would not have been possible without the outstanding work of the whole team at SEDDI, which has played roles in important tasks, including administrative work, dataset creation, financial support, help with software and infrastructure, as well as valuable scientific discussions. My individual contributions to each project can be inferred from my relative position in the author list.</p>
</sec>
</sec>
<sec id="c1-s4">
<label>1.4</label>
<title>Measurable Outcomes</title>
<sec id="c1-s4.s1">
<label>1.4.1</label>
<title>Publications &#x0026; Patents</title>
<sec id="c1-s4.s1.1">
<title>Peer Reviewed Publications</title>
<p>The contributions in this thesis have led to the following publications, listed in chronological order:</p>
<list list-type="bullet">
<list-item><p><italic>&#x201C;Neural Photometry-guided Visual Attribute Transfer&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Elena Garces; <italic>IEEE Transactions on Visualization and Computer Graphics (TVCG), 2021</italic> [<xref ref-type="bibr" rid="CIT338">RG21</xref>].</p>
<p>This journal has an impact factor of 5.226, and its position in the JCR index is 13 out of 110 (Q1) in the category Computer Science, Software Engineering (data from 2021).</p></list-item>
<list-item><p><italic>&#x201C;A Survey on Intrinsic Images: Delving Deep into Lambert and Beyond&#x201D;</italic></p>
<p>Elena Garces, Carlos Rodriguez-Pardo, Dan Casas, Jorge Lopez-Moreno; <italic>International Journal in Computer Vision (IJCV), 2022</italic> [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>].</p>
<p>This journal has an impact factor of 13.369, and its position in the JCR index is 10 out of</p>
<p>145 (Q1) in the category Computer Science, Artificial Intelligence (data from 2021).</p></list-item>
<list-item><p><italic>&#x201C;SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Elena Garces; <italic>IEEE Transactions on Visualization and Computer Graphics (TVCG), 2022</italic> [<xref ref-type="bibr" rid="CIT339">RG22</xref>].</p>
<p>This journal has an impact factor of 5.226, and its position in the JCR index is 13 out of</p>
<p>110 (Q1) in the category Computer Science, Software Engineering (data from 2021)</p></list-item>
<list-item><p><italic>&#x201C;How Will It Drape Like? Capturing Fabric Mechanics From Depth Images&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Melania Prieto-Martin, Dan Casas, Elena Garces; <italic>Computer Graphics Forum (Proceedings of Eurographics 2023)</italic> [<xref ref-type="bibr" rid="CIT340">Rod+23b</xref>].</p>
<p>This journal has an impact factor of 2.363, and its position in the JCR index is 53 out of</p>
<p>110 (Q2) in the category Computer Science, Software Engineering (data from 2021).</p></list-item>
<list-item><p><target target-type="page" id="pges_24"/><italic>"UMat: Uncertainty-Aware Single Image High Resolution Material Capture"</italic></p>
<p>Carlos Rodriguez-Pardo, Henar Dominguez, David Pascual, Elena Garces; <italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023</italic> [<xref ref-type="bibr" rid="CIT337">Rod+23a</xref>].</p>
<p>This conference is the highest impact venue in computer science according to <italic>Google Scholar</italic>, and is rated as a CORE <italic>A++</italic> conference (data from 2021).</p></list-item>
<list-item><p><italic>&#x201C;Towards Material Digitization with a Dual-scale Optical System&#x201D;</italic></p>
<p>Elena Garces, Victor Arellano, Carlos Rodriguez-Pardo, David Pascual, Sergio Suja, Jorge Lopez-Moreno; <italic>ACM Transactions on Graphics (Proceedings of SIGGRAPH), 2023</italic> [<xref ref-type="bibr" rid="CIT103">Gar+23</xref>]. This journal has an impact factor of 7.403, and its position in the JCR index is 9 out of 110 (Q1) in the category Computer Science, Software Engineering (data from 2021).</p></list-item>
<list-item><p><italic>&#x201C;NeuBTF: Neural Fields for BTF Encoding and Transfer&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Konstantinos Kazatzis, Jorge Lopez-Moreno, Elena Garces; 2023</p>
<p><italic>This paper is currently under peer review.</italic></p></list-item>
<list-item><p><italic>&#x201C;NEnv: Neural Environment Maps for Global Illumination&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Javier Fabre, Elena Garces, Jorge Lopez-Moreno; 2023</p>
<p><italic>This paper is currently under peer review.</italic></p></list-item>
</list>
</sec>
<sec id="c1-s4.s1.2">
<title>Patents</title>
<p>Some of the work developed during my PhD has resulted in two patent applications:</p>
<list list-type="bullet">
<list-item><p><italic>&#x201C;Generation of macro-scale material property maps from images and micro-scale properties&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Elena Garces; <italic>(SEDDI 2020, Filed Patent Status)</italic></p></list-item>
<list-item><p><italic>&#x201C;Neural synthesis of tileable textures&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Elena Garces; <italic>(SEDDI 2020, Filed Patent Status)</italic></p></list-item>
</list>
</sec>
</sec>
<sec id="c1-s4.s2">
<label>1.4.2</label>
<title>Awards &#x0026; Other Merits</title>
<p>We include here a list of awards and other merits received throughout the development of this thesis.</p>
<list list-type="bullet">
<list-item><p><italic>Top Reviewer</italic> recognition at the 2022 Conference on Neural Information Processing Systems (NeurIPS).</p></list-item>
<list-item><p> <italic>Outstanding Reviewer</italic> recognition at the 2022 European Conference on Computer Vision (ECCV).</p></list-item>
<list-item><p>Our work on <italic>Neural Visual Attribute Transfer</italic> [<xref ref-type="bibr" rid="CIT338">RG21</xref>] (<xref ref-type="book-part" rid="c2">Chapter 2</xref>) was invited to the <italic>ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games</italic>.</p></list-item>
<list-item><p><target target-type="page" id="pges_25"/>Our work on Generative Models for Tileable Texture Synthesis [<xref ref-type="bibr" rid="CIT339">RG22</xref>] (<xref ref-type="book-part" rid="c3">Chapter 3</xref>) was presented at <italic>Machine Learning Tokyo</italic> as an invited talk, and as a poster presentation at the AI for Content Creation Workshop (AI4CC) at CVPR 2022.</p></list-item>
</list>
</sec>
<sec id="c1-s4.s3">
<label>1.4.3</label>
<title>Student Supervision</title>
<p>During the development of this thesis, I have co-supervised the bachelor&#x2019;s final thesis of four computer science students:</p>
<list list-type="bullet">
<list-item><p><italic>&#x201C;Multimodal Exploration of Textile Databases Based on Perceptual Parameters&#x201D;</italic></p>
<p>Mar&#x00ED;a Pilar Alcarria Peinado, 2021. <italic>Final grade: 94/100</italic></p></list-item>
<list-item><p><italic>&#x201C;A Fashion Recommender System Based on Multimodal Attributes&#x201D;</italic></p>
<p>Gonzalo Llosa, 2023. <italic>(Ongoing)</italic></p></list-item>
<list-item><p><italic>&#x201C;Image Generative Models for Garment Design&#x201D;</italic></p>
<p>Alfonso Chiclana, 2023. <italic>(Ongoing)</italic></p></list-item>
<list-item><p><italic>&#x201C;Web Interfaces for the Exploration of Text-to-Image Generative Models&#x201D;</italic></p>
<p>Carlos Cepeda, 2023. <italic>(Ongoing)</italic></p></list-item>
</list>
</sec>
<sec id="c1-s4.s4">
<label>1.4.4</label>
<title>Teaching, Talks, and Seminars</title>
<p>During the first half of 2023, I worked as an adjunct lecturer at Universidad Carlos III in Madrid, where I taught and led a mandatory graduate course on <italic>Intelligent Data Analysis</italic>, at the MSc in Computer Engineering.</p>
<p>Throughout my Ph.D., I had the opportunity to disseminate my work across different formal and informal events. I presented some of the work in this Ph.D. thesis on different venues, including:</p>
<list list-type="bullet">
<list-item><p><italic>ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games (I3D), 2022</italic>: I had the chance to present the work in <xref ref-type="book-part" rid="c2">Chapter 2</xref> on propagating spatially-varying attributes of materials, as an invited talk of our TVCG paper.</p></list-item>
<list-item><p><italic>Machine Learning Tokyo init, May 2022</italic>: I presented the work in <xref ref-type="book-part" rid="c3">Chapter 3</xref> on tileable texture synthesis as an invited seminar.</p></list-item>
<list-item><p><italic>AI4CC workshop at CVPR, June 2022; &#x0026; ELLIS Doctoral Consortium, September 2022</italic>: I presented the work in <xref ref-type="book-part" rid="c3">Chapter 3</xref> on tileable texture synthesis as poster presentations.</p></list-item>
<list-item><p><italic>EUROGRAPHICS, May 2023</italic>: We will present the work in <xref ref-type="book-part" rid="c6">Chapter 6</xref> on casual and perceptually-validate mechanical parameter capture of fabrics.</p></list-item>
<list-item><p><target target-type="page" id="pges_26"/><italic>CVPR, June 2023</italic>: We will present the work in <xref ref-type="book-part" rid="c5">Chapter 5</xref> on uncertainty-aware high resolution SVBRDF digitization.</p></list-item>
</list>
</sec>
<sec id="c1-s4.s5">
<label>1.4.5</label>
<title>Reviewing Activity</title>
<p>During the development of this thesis, I have served as a reviewer at multiple conferences, workshops, and journals on computer vision and machine learning:</p>
<list list-type="bullet">
<list-item><p><italic>British Machine Vision Conference (BMVC)</italic>: 2020</p></list-item>
<list-item><p><italic>European Conference on Computer Vision (ECCV)</italic>: 2022, Outstanding Reviewer</p></list-item>
<list-item><p><italic>IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic>: 2020, 2022, 2023</p></list-item>
<list-item><p><italic>IEEE International Conference on Computer Vision (ICCV)</italic>: 2021, 2023</p></list-item>
<list-item><p><italic>International Conference on Learning Representations (ICLR)</italic>: 2022</p></list-item>
<list-item><p><italic>International Conference on Machine Learning (ICML)</italic>: 2023</p></list-item>
<list-item><p><italic>LatinX in Computer Vision Research Workshop</italic>: 2023</p></list-item>
<list-item><p><italic>Neural Information Processing Systems (NeurIPS)</italic>: 2022 Top Reviewer, 2023</p></list-item>
<list-item><p><italic>The Visual Computer Journal (TVCJ)</italic>: 2020-2022</p></list-item>
<list-item><p><italic>Women in Computer Vision Workshop (WiCV)</italic> : 2022, 2023</p></list-item>
</list>
</sec>
<sec id="c1-s4.s6">
<label>1.4.6</label>
<title>Industrial Research Projects and Commercial Products</title>
<p>This thesis has been developed within an industrial Ph.D. in SEDDI, a startup based in Madrid, Spain, which has the goal of providing deep tech solutions to the textile and fashion industries. A significant part of the work presented in this thesis is part of commercially-available products and industry research projects. Notably, some of the algorithms (or variations thereof) presented in <xref ref-type="book-part" rid="c2">Chapters 2</xref>, <xref ref-type="sec" rid="c2-sC">2.C</xref> and <xref ref-type="book-part" rid="c5">5</xref> are part of <uri xlink:href="http://www.Textura.ai">Textura.ai</uri>, an AI-powered commercially available material digitization product. Further, the <italic>drape similarity metric</italic> presented in <xref ref-type="book-part" rid="c6">Chapter 6</xref> has also been a part of a project developed for a textile manufacturer.</p>
</sec>
</sec>
</body>
</book-part>
<book-part id="c2" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_27"/>Chapter 2.</label>
<title>Visual Attribute Transfer with Photometry-guided Neural Networks</title>
</title-group>
</book-part-meta>
<body>
<fig id="fig-2.1">
<label>Figure 2.1:</label>
<caption><title>Illustration of the capabilities of our attribute transfer framework. On the left, we show a subset of our <italic>photometric dataset</italic>, comprised of several images of a material captured with different illumination conditions. On their right, two visual attributes: surface normals and segmentation maps. Using a larger image of the same or similar materials as <italic>guidance</italic>, our method can propagate the visual attributes, achieving high resolution, robust and controllable transfers at small computational cost.</title></caption>
<alt-text>Illustration of the capabilities of our attribute transfer framework. On the left, we show a subset of our photometric dataset, comprised of several images of a material captured with different illumination conditions. On their right, two visual attributes: surface normals and segmentation maps. Using a larger image of the same or similar materials as guidance, our method can propagate the visual attributes, achieving high resolution, robust and controllable transfers at small computational cost.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.1.jpg"/>
</fig>
<p>This chapter presents a deep learning-based method capable of propagating spatially-varying visual attributes of a material (e.g. texture maps or stylizations) to larger samples of the same or similar materials. By leveraging the material&#x2019;s photometric response and a comprehensive data augmentation policy, we make the transfer robust to novel illumination conditions and affine deformations of the underlying surface. Our model relies on a supervised image-to-image translation framework and is agnostic to the transferred visual domain; we showcase a semantic segmentation, a normal map, and a stylization. Following an image analogies approach, the method only requires the training data to contain the same visual structures as the input guidance. Our approach works at interactive rates, making it suitable for material edit applications. We thoroughly evaluate our learning methodology in a controlled setup providing quantitative measures of performance. Last, we demonstrate that training the model on a single material is enough to generalize to materials of the same type without the need for massive datasets. The contributions presented in this chapter have led to the following publication<sup>1</sup>, which was also presented at the <italic>2022 ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games (I3D 2022)</italic> as an invited talk:</p>
<fig id="fig-1">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-1.jpg"/>
</fig>
<disp-quote>
<p><target target-type="page" id="pges_28"/>&#x201C;Neural Photometry-guided Visual Attribute Transfer&#x201D;</p>
<p>Carlos Rodriguez-Pardo, Elena Garces</p>
<p><italic>IEEE Transactions on Visualization and Computer Graphics (TVCG)</italic></p>
<p>(2021)</p>
</disp-quote>
<sec id="c2-s1">
<label>2.1</label>
<title>Introduction</title>
<fig id="fig-2.2">
<label>Figure 2.2:</label>
<caption><title>The input exemplars on the left-bottom are transferred to the input guidances (second column) using images of the material taken under multiple illuminations (photometric input) as training data. Using this data for training, Texler <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>] fail to generalize to geometric distortions and illumination conditions not present in the training set.</title></caption>
<alt-text>The input exemplars on the left-bottom are transferred to the input guidances (second column) using images of the material taken under multiple illuminations (photometric input) as training data. Using this data for training, Texler et al. [Tex+20b] fail to generalize to geometric distortions and illumination conditions not present in the training set.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.2.jpg"/>
</fig>
<p>The development of effective and editable material models is becoming increasingly important so that users of virtual prototyping, video-games or AR/VR applications can have compelling and realistic experiences. For instance, allowing for the creation of virtual environments that surpass the uncanny valley, or empowering artists to create breathtaking visual settings. Effective material representations require understanding the properties that uniquely define them, which we refer to as <italic>visual material attributes</italic>. These attributes are spatially-varying parameters that maintain spatial coherency with respect to the material structure while remaining invariant to changes in the scene illumination or the geometry of the <target target-type="page" id="pges_29"/>underlying object. For example, they may represent optical properties of a microfacet spatially-varying BRDF (albedo, normals, roughness, anisotropy, etc.), but also artistic stylizations or higher-level properties, as in semantic segmentation masks. The former can be obtained with specific optical capture devices and estimation methods, either using machine learning [<xref ref-type="bibr" rid="CIT271">Mer+19</xref>] or classic optimization methods [<xref ref-type="bibr" rid="CIT374">Ste+14</xref>]. The latter may require time-consuming human labor.</p>
<p>Obtaining these attributes for large material samples can be problematic in certain situations. For instance, measuring high-resolution surface normals of large material samples or generating densely annotated segmentation masks, require either unfeasibly large capturing setups or expensive manual input. A common strategy to handle these cases consists of propagating these attributes using a larger image of the same material used as guidance image. This problem has been formulated within the context of image analogies [<xref ref-type="bibr" rid="CIT150">Her+01</xref>]. Propagating, or transferring, these properties requires preserving the local spatial regularities of the material, as well as adapting to the global variations of the guidance image, which is a highly challenging problem.</p>
<p>Previous work has used PatchMatch-based synthesis [<xref ref-type="bibr" rid="CIT269">Mel+12</xref>], look-up-tables [<xref ref-type="bibr" rid="CIT336">RPG16</xref>], or neural networks [<xref ref-type="bibr" rid="CIT265">Maz+19</xref>] by looking for repetitive patterns, match color statistics, or image gradients. These methods are however prohibitive due to their runtime performance, taking hours. Image analogies methods have been extended so as to leverage the power of convolutional neural networks as image descriptors [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>; <xref ref-type="bibr" rid="CIT025">Ben+21</xref>], but those methods do not yield predictable outputs, and their computational complexity yields them impractical in interactive applications. Recently, neural networks have proven successful for the task of video stylization [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>], where the user needs to provide a few input editing exemplars which are propagated at interactive rates to the rest of the video. We show that this approach lacks the generalization power to work with images taken under unseen resolutions and geometric or illumination conditions.</p>
<p>In this chapter, we propose a novel learning-based method to propagate any visual material attribute &#x015B;estimated locally for a material&#x015B; to larger samples of it. We train a neural network per material using image-to-image translation methods, making use of a policy of data augmentation that makes the transfer invariant to affine transformations (scale, rotations and shears). As opposed to other methods which blindly apply data augmentation strategies, our method is robustly evaluated using a real dataset to guarantee predictability in the estimations. Further, illumination invariance is obtained by feeding the network with multiple images of the material taken under a diverse set of illuminations. Our models can be trained in less than a minute and generalize to materials with the same microestructure.</p>
<p><target target-type="page" id="pges_30"/>In summary, we present the following contributions:</p>
<list list-type="bullet">
<list-item><p>The first method to use photometric data to train an image-to-image translation model capable of propagating any kind of visual material attribute to larger samples of the material, regardless of the illumination conditions of the input image.</p></list-item>
<list-item><p>A data augmentation policy, thoroughly evaluated with a real dataset, designed to make the transfer invariant to affine deformations.</p></list-item>
<list-item><p>Exhaustive comparisons with related work, demonstrating that we can achieve more predictable and higher-quality mappings with a fraction of the computational cost.</p></list-item>
<list-item><p>We show that our trained models generalize to materials with similar microstructure as the ones used for training.</p></list-item>
<list-item><p>On <xref ref-type="sec" rid="c2-sC">Section 2.C</xref>, we show an extension on this method for High Resolution SVBRDF capture, which is part of a publication accepted to <italic>SIGGRAPH 2023</italic>.</p></list-item>
</list>
</sec>
<sec id="c2-s2">
<label>2.2</label>
<title>Related Work</title>
<p>Several computer graphics and vision problems are closely related to our method. The most similar ones are those related to any form of by-example <italic>visual attribute transfer</italic> (e.g color, texture, style, or geometry). Besides, we also review material estimation and capture methods.</p>
<sec id="c2-s2.s1">
<label>2.2.1</label>
<title>Visual Attribute Transfer</title>
<p>Visual attribute transfer refers to the problem of transferring some visual attributes (e.g. color, style, texture, or geometry) of one or many exemplars to another exemplar while preserving its <italic>content</italic>.</p>
<p>This problem can be formulated within the context of <italic>Image Analogies</italic> [<xref ref-type="bibr" rid="CIT150">Her+01</xref>], in which the goal is to stylize a target un-stylized image <italic>B</italic>, given a pair of images <italic>A</italic> (un-stylized) and <italic>A&#x2019;</italic> (stylized). The most common approach to tackle this problem has been via patch-based texture synthesis [<xref ref-type="bibr" rid="CIT026">B&#x00E9;n+13</xref>; <xref ref-type="bibr" rid="CIT023">Bar+15</xref>; <xref ref-type="bibr" rid="CIT178">Jam+19</xref>]. More recent approaches have leveraged the capabilities of deep latent spaces within convolutional neural networks to disentangle style from content [<xref ref-type="bibr" rid="CIT326">Ree+15</xref>]. A seminal work by Gatys <italic>et al.</italic>[<xref ref-type="bibr" rid="CIT108">GEB15b</xref>] utilizes a VGG-19 convolu-tional neural network [<xref ref-type="bibr" rid="CIT364">SZ15</xref>] pre-trained on ImageNet [<xref ref-type="bibr" rid="CIT064">Den+09</xref>] as a feature descriptor for images, in which style and content are associated with different layers of the network, and transferred by gradient descent optimization. Their work on <italic>style transfer</italic> has been extended for single images [<xref ref-type="bibr" rid="CIT168">HB17</xref>; <xref ref-type="bibr" rid="CIT238">Li+17b</xref>; <xref ref-type="bibr" rid="CIT186">JAF16</xref>; <xref ref-type="bibr" rid="CIT046">Che+17b</xref>] and video [<xref ref-type="bibr" rid="CIT045">Che+17a</xref>], as well as for developing image-space distance metrics which resemble human perception [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>]. A shared limitation of many of these <target target-type="page" id="pges_31"/>methods is their narrow capabilities to provide predictable edits, and considerable focus nowadays is put towards this end [<xref ref-type="bibr" rid="CIT110">Gat+17</xref>; <xref ref-type="bibr" rid="CIT120">Gu+18</xref>]. A comprehensive review on the topic of neural style transfer is provided by Jing et al. [<xref ref-type="bibr" rid="CIT185">Jin+19</xref>]. Our work differs from traditional style transfer approaches in the sense that we deal with a more constrained problem that requires predictable outcomes.</p>
<p>Exploiting the power of deep neural networks in the image analogies problem was tackled by Liao <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>] who, by assuming a semantic prior over an exemplar input image and a target one, propose a method capable of finding a bijective mapping between both inputs, enabling two-way stylizations. Single-image generative models [<xref ref-type="bibr" rid="CIT359">SDM19</xref>] were extended to the image analogies problem in [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>], by using convolutional neural networks to generate a new image with the style of an input <italic>style</italic> and the <italic>structure</italic> of another image. These methods, however, rely heavily on content or semantic features, making them vulnerable to lighting or geometric differences between the input images; and are computationally expensive, rendering them impractical for interactive applications. Our method is robust to both geometric distortions and illumination variations, and works at interactive rates. Similar in spirit to our method, as it explicitly considers texture variations due to illumination, is the work of Fi&#x0161;er <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT093">Fi&#x0161;+16</xref>], which applies patch-match to provide illumination-dependent exemplar-based stylizations to cartoon pictures. In contrast, our work is meant to be illumination-invariant. Also concerned with stylization problems, Texler <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>] present a method highly related to ours. They apply patch-based training of an encoder-decoder deep neural network using key frames stylized by a user. Resembling our approach, their algorithm also follows a few-shot learning strategy using as training data a few exemplar patches. However, as opposed to our method, they do not account for the variability of the appearance of materials under different lighting and viewpoint so, as we show in <xref ref-type="sec" rid="c2-s7">Section 2.7</xref>, their method does not generalize to unseen illuminations or geometric variations.</p>
<p><italic>Image colorization</italic> is concerned with colorizing a gray-scale image given a few colorized exemplars. In this problem, it is critical to infer semantic relationships between the images so that the new scene is perceptually coherent and plausible [<xref ref-type="bibr" rid="CIT140">He+18</xref>; <xref ref-type="bibr" rid="CIT442">Zha+19a</xref>; <xref ref-type="bibr" rid="CIT141">He+19</xref>]. Similarly, <italic>edit propagation</italic> methods [<xref ref-type="bibr" rid="CIT010">AP08</xref>; <xref ref-type="bibr" rid="CIT087">End+16</xref>] work by propagating strokes provided by a user to the rest of the image, removing the need for a semantic understanding of the input image. Our work is related to the latter techniques, as we leverage the feature spaces of the CNNs and do not require a large labeled dataset to effectively solve our problem, and it can also be used to propagate segmentation masks [<xref ref-type="bibr" rid="CIT239">LAA08</xref>].</p>
</sec>
<sec id="c2-s2.s2">
<label>2.2.2</label>
<title>Textured Materials</title>
<p>Many real-world materials show spatial regularities, commonly referred to as <italic>textures</italic>. The patterns present in textures can be parameterized, which <target target-type="page" id="pges_32"/>allows for low-cost material capture or synthesis models. A way to model textured materials is through BTFs (Bidirectional Texture Functions) [<xref ref-type="bibr" rid="CIT058">Dan+99</xref>; <xref ref-type="bibr" rid="CIT233">LM01</xref>], a technique that uses multiple camera views and lighting angles to capture a dense sampling of the appearance of a material. Inspired by such methods, we leverage several images of the material under different illumination conditions, however, we require fewer data than typical BTFs capture setups [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>; <xref ref-type="bibr" rid="CIT316">Rai+20</xref>]. Similarly, the problem of extrapolating BTFs captures to larger material samples was addressed by Steinhausen et al. [<xref ref-type="bibr" rid="CIT375">Ste+15a</xref>; <xref ref-type="bibr" rid="CIT376">Ste+15b</xref>], who propagate measured BTFs using texture synthesis. Our method is not meant to propagate full BTFs measurements but could potentially be applied to such datasets.</p>
<p>The goal of texture synthesis is to reconstruct a larger image given a small sample leveraging structural content. This is a long-standing problem in the computer graphics field and different strategies have been proposed, for instance, using PatchMatch [<xref ref-type="bibr" rid="CIT070">Dia+15</xref>], texture transport [<xref ref-type="bibr" rid="CIT005">AWL15</xref>], point processes [<xref ref-type="bibr" rid="CIT123">Gue+20</xref>; <xref ref-type="bibr" rid="CIT231">LH06</xref>], or neural networks [<xref ref-type="bibr" rid="CIT085">EM17</xref>; <xref ref-type="bibr" rid="CIT458">Zho+18</xref>; <xref ref-type="bibr" rid="CIT095">FAW19</xref>; <xref ref-type="bibr" rid="CIT341">Rod+19</xref>]. Also related to our work, Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT247">LPG19</xref>] capture the appearance of materials by first estimating their BRDF and, then, synthesizing the high resolution microstructure from a dataset of measured SVBRDFs. Our problem is unlike texture synthesis, as we do not aim to create novel content but to predictably transfer visual material attributes.</p>
</sec>
<sec id="c2-s2.s3">
<label>2.2.3</label>
<title>SVBRDF Estimation</title>
<p>The problem of estimating a SVBRDF model from one or several images using lightweight capture setups is becoming increasingly popular in the literature. Early work [<xref ref-type="bibr" rid="CIT151">HS05</xref>; <xref ref-type="bibr" rid="CIT152">HS03</xref>] leveraged Photometric Stereo [<xref ref-type="bibr" rid="CIT174">Ike81</xref>] and SVBRDF manifold bootstrapping [<xref ref-type="bibr" rid="CIT077">Don+10</xref>] for surface geometry reconstruction, while newer methods exploit the power of deep neural networks. Recent surveys by Guarnera [<xref ref-type="bibr" rid="CIT121">Gua+16</xref>] and Dong [<xref ref-type="bibr" rid="CIT076">Don19</xref>] contain relevant approaches. While our method is not meant to estimate the SVBRDF properties of a material, it can be used in combination with those techniques to create larger material assets.</p>
<p>There are a few methods that follow a similar paradigm to ours, transferring pre-estimated SVBRDF maps to a larger material sample. Using PatchMatch texture synthesis, Melendez <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT269">Mel+12</xref>] transfer displacement and albedo maps from small samples of the materials. Their method is limited to daylight illumination and materials present in fa&#x00E7;ades. By means of look-up-tables, and using surface normals and specular as guidance, Riviere <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT336">RPG16</xref>] transfer surface reflectance captured with controlled LCD lighting to a material sample observed under natural lighting. Recently, Deschaintre <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT066">Des+19</xref>] fine-tune a network trained to estimate SVBRDFs [<xref ref-type="bibr" rid="CIT065">Des+18</xref>], to work on larger material samples taking a guidance image as input. This approach is limited <target target-type="page" id="pges_33"/>to transfer a pre-defined set of property maps while our method can transfer any kind. The strategy of using multiple images of the material under different illuminations as input data is not new. Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT235">Li+17a</xref>] and Ye <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT437">Ye+18</xref>] utilize a self-augmentation strategy to make the estimation of the SVBRDF more robust to unknown environment illumination. We are inspired by these approaches to increase robustness in the model predictions.</p>
</sec>
</sec>
<sec id="c2-s3">
<label>2.3</label>
<title>Problem Formulation</title>
<p>Our goal is to transfer a <italic>D</italic>-dimensional spatially-varying <italic>visual attribute w</italic> of a material (for example, estimated locally at high resolution) to a larger sample of it.</p>
<p>We formulate the problem with an image-to-image translation approach. For training, our method takes as input: a <italic>photometric</italic> dataset <inline-formula><mml:math id="M54" display='block'><mml:msub><mml:mi mathvariant="script">I</mml:mi><mml:mi mathvariant="normal">L</mml:mi></mml:msub></mml:math></inline-formula>, and a <italic>visual attribute w</italic>. The <italic>photometric</italic> dataset consisting of a number of RGB planar images of a material, <inline-formula><mml:math id="M55" display='block'><mml:msub><mml:mi mathvariant="script">I</mml:mi><mml:mi mathvariant="normal">L</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="script">I</mml:mi><mml:mi mathvariant="normal">L</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, illuminated with different light sources <inline-formula><mml:math id="M56" display='block'><mml:mi mathvariant="bold-italic">l</mml:mi></mml:math></inline-formula> &#x2208; L, of <italic>n</italic> &#x00D7; <italic>m</italic> pixels. This kind of images can be either captured with specific devices [<xref ref-type="bibr" rid="CIT291">Nam+16</xref>; <xref ref-type="bibr" rid="CIT272">MWK17</xref>; <xref ref-type="bibr" rid="CIT007">Alc+19</xref>], or synthetically rendered given an inverse material estimation pipeline [<xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT067">DDB20</xref>]. The visual attribute being a spatially-varying map <italic>w</italic> &#x2208; &#x211D;<sup><italic>n</italic>&#x00D7;<italic>m</italic>&#x00D7;<italic>D</italic></sup> of any kind, and dimensions <italic>D</italic>, that maintains pixel-wise correspondence with the photometric input images <inline-formula><mml:math id="M57" display='block'><mml:msub><mml:mi mathvariant="script">I</mml:mi><mml:mi mathvariant="normal">L</mml:mi></mml:msub></mml:math></inline-formula>. <xref ref-type="fig" rid="fig-2.3">Figure 2.3</xref> and <xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref> show examples of these images for three different visual attributes: a stylization, a segmentation, and a normal map.</p>
<fig id="fig-2.3">
<label>Figure 2.3:</label>
<caption><title>Overview of the method. We learn a mapping between the photometric response <italic>I<sub>L</sub></italic> of the material and a visual property map <italic>w</italic>. We make this mapping robust to affine transformations by means of a particular policy of data augmentation used for training. We learn one model <inline-formula><mml:math id="M51" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> per material and visual attribute, which allows us to robustly evaluate the performance of the method under several transformations of the guidance images X. At evaluation time, <inline-formula><mml:math id="M52" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> can have any size. The training of <inline-formula><mml:math id="M53" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> per visual attribute <italic>w</italic> takes less than a minute.</title></caption>
<alt-text>Overview of the method. We learn a mapping between the photometric response IL of the material and a visual property map w. We make this mapping robust to affine transformations by means of a particular policy of data augmentation used for training. We learn one model M per material and visual attribute, which allows us to robustly evaluate the performance of the method under several transformations of the guidance images X. At evaluation time, M can have any size. The training of M per visual attribute w takes less than a minute.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.3.jpg"/>
</fig>
<fig id="fig-2.4">
<label>Figure 2.4:</label>
<caption><title>The color augmentation policy takes advantage of the material structural regularities to make the transfer robust to different albedos. (a) A diffuse image of the portion of the material used for training, and its corresponding normal map. (b) Input guidance image. (c) Transferred normals without the color augmentation policy. (d) Transferred normals using color augmentation. Note that the training data does not include images containing the white yarn.</title></caption>
<alt-text>The color augmentation policy takes advantage of the material structural regularities to make the transfer robust to different albedos. (a) A diffuse image of the portion of the material used for training, and its corresponding normal map. (b) Input guidance image. (c) Transferred normals without the color augmentation policy. (d) Transferred normals using color augmentation. Note that the training data does not include images containing the white yarn.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.4.jpg"/>
</fig>
<fig id="fig-2.5">
<label>Figure 2.5:</label>
<caption><title>An overview of our evaluation dataset. (a) Five example images of our high resolution captured data illuminated with diffuse light and four directional light sources. (b) Examples of some of our <italic>visual property maps</italic>: colorization, normals and yarn segmentation. (c) Evaluation images taken under diffuse and directional illumination sources. <italic>Denim</italic>, (c) contains labels for several patches used for the evaluation in <xref ref-type="sec" rid="c2-s6">Section 2.6</xref>. Those materials show different properties that may prove challenging for our visual attribute transfer task. Our <italic>denim</italic> and <italic>knit</italic> fabrics show diverse color variations (different colored dyed yarns and stochastic albedo, respectively), <italic>linen</italic> shows strong geometric variations and <italic>satin</italic> shows anisotropic optical behavior.</title></caption>
<alt-text>An overview of our evaluation dataset. (a) Five example images of our high resolution captured data illuminated with diffuse light and four directional light sources. (b) Examples of some of our visual property maps: colorization, normals and yarn segmentation. (c) Evaluation images taken under diffuse and directional illumination sources. Denim, (c) contains labels for several patches used for the evaluation in Section 2.6. Those materials show different properties that may prove challenging for our visual attribute transfer task. Our denim and knit fabrics show diverse color variations (different colored dyed yarns and stochastic albedo, respectively), linen shows strong geometric variations and satin shows anisotropic optical behavior.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.5.jpg"/>
</fig>
<p>Given this data for a single material, and a strategy of patch-based training, we learn a function <inline-formula><mml:math id="M58" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> that can be applied to a new <italic>guidance</italic> image of the input material (or a similar one) <inline-formula><mml:math id="M59" display='block'><mml:mi mathvariant="normal">X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> of any size <italic>N</italic> &#x00D7; <italic>M</italic>, to get its corresponding visual attribute <inline-formula><mml:math id="M60" display='block'><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>:</p>
<p><disp-formula id="Eq002-1"><label><target target-type="page" id="pges_34"/>(2.1)</label> <mml:math id="M61" display='block'><mml:mi mathvariant="script">M</mml:mi><mml:mo>:</mml:mo><mml:mi mathvariant="normal">X</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><disp-formula id="Eq002-2"><label>(2.2)</label> <mml:math id="M62" display='block'><mml:mtext>s.t.</mml:mtext><mml:msub><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mover><mml:mrow/><mml:mpadded><mml:mi mathvariant="script">M</mml:mi><mml:mspace/></mml:mpadded></mml:mover><mml:mi>&#x03C9;</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:msub><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi mathvariant="script">I</mml:mi><mml:mi mathvariant="normal">L</mml:mi></mml:msub></mml:math></disp-formula></p>
<p>At evaluation time, the guidance image X might contain different colors, scales, illumination, or affine distortions than the images used for training. <xref ref-type="fig" rid="fig-2.3">Figure 2.3</xref> shows an overview of the training and evaluation processes. In <xref ref-type="sec" rid="c2-s6">Section 2.6</xref>, we evaluate the conditions of the input guidance image upon which the method provides robust estimations.</p>
</sec>
<sec id="c2-s4">
<label>2.4</label>
<title>Learning Framework</title>
<p>In this section, we describe the patch-based and data augmentation strategies used for training, the neural network design and loss functions, and the implementation details.</p>
<sec id="c2-s4.s1">
<label>2.4.1</label>
<title>Patch-based Training</title>
<p>Each training step takes as input a pair of corresponding image patches taken from the photometric input <inline-formula><mml:math id="M63" display='block'><mml:msub><mml:mi mathvariant="script">I</mml:mi><mml:mi mathvariant="normal">L</mml:mi></mml:msub></mml:math></inline-formula>, and the visual attribute <italic>w</italic>. Using this data alone already provides a good starting point for generalizing to unseen illumination setups. However, it is not sufficient in scenarios in which the guidance image contains variations due to image noise, a different scale, or any other affine distortion. In order to make the transfer invariant to these transformations, the network needs to be trained with the appropriate data. Data augmentation strategies are essential for reducing the amount of necessary data for training <target target-type="page" id="pges_35"/>[<xref ref-type="bibr" rid="CIT363">SK19</xref>; <xref ref-type="bibr" rid="CIT194">Kar+20a</xref>], however, random strategies not taking into account the particular domain might degrade the quality of the prediction. We therefore follow a pre-defined data augmentation policy <inline-formula><mml:math id="M64" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> (illustrated in <xref ref-type="fig" rid="fig-2.3">Figure 2.3</xref>), where random operations are performed sequentially.</p>
<p><bold>Color Augmentation</bold> Even if the microstructure of the material is homogeneous and can be measured using only a small patch of it, it may not be possible for the network to estimate its visual attributes for parts of the material which contain previously unseen colors. We correct this by randomly permuting the color channels of the photometric input (see <xref ref-type="fig" rid="fig-2.4">Figure 2.4</xref>). As we show in <xref ref-type="sec" rid="c2-s7">Section 2.7</xref>, this data augmentation policy also helps make the model generalize the transfer to similar materials, by learning features that are more related to the structure of the material than to its color. Similar operations have been recently proposed for finding robust visual representations on self-supervised settings [<xref ref-type="bibr" rid="CIT047">Che+20b</xref>]. Only the photometric input is subject to this transformation.</p>
<p><bold>Affine Transforms</bold> In order to allow for material editing applications in real images, the model should generalize to images taken under camera perspectives and geometry distortions different than those present in the fronto-planar images used for training. Many of such texture irregularities can be defined as an affine transform: translation, rotation, shear, or scaling. As CNNs are shift-invariant by design [<xref ref-type="bibr" rid="CIT200">Kau17</xref>], we propose to augment our datasets with random transformations for scalings, rotations, and shears. Those transforms can be efficiently performed on images through matrix multiplications. However, some visual property maps need to be treated specially, as the spatial transform might have a different behavior in 3<italic>D</italic> vector space, e.g. normals or tangents. In those maps, each pixel is a representation of a 3<italic>D</italic> vector. As such, we also perform the rotations and shear operations to these maps in 3<italic>D</italic> space, by multiplying each normal vector by the same affine transformation matrix applied to the 2<italic>D</italic> image. We perform the rotation around the <italic>Z</italic> axis, thus, assuming the camera sensor is parallel to the object plane.</p>
<p><bold>Cropping</bold> Inspired by recent work on patch-based learning [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>; <xref ref-type="bibr" rid="CIT303">Par+20</xref>], during training, the network receives small patches of each input pair. Those patches are randomly cropped from the randomly augmented images, so the network receives a considerable amount of variations of the same material, thus making generalization possible.</p>
</sec>
<sec id="c2-s4.s2">
<label>2.4.2</label>
<title>Network Design</title>
<p>We follow a uni-modal image-to-image translation learning strategy, assuming there is only one correct mapping from input to output image. This approach is reasonable for the kind of transfers we test in this work. However, it might <target target-type="page" id="pges_36"/>fail for ambiguous cases where there are multiple suitable outputs for the same input, for which multi-modal approaches [<xref ref-type="bibr" rid="CIT461">Zhu+17b</xref>] or GANs are more advisable, albeit harder to train. Specifically, our model <inline-formula><mml:math id="M65" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> is a shallow U-Net network [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] with 4 blocks of layers, containing a small number of trainable parameters, inspired on few-shot learning strategies [<xref ref-type="bibr" rid="CIT418">Wan+20d</xref>; <xref ref-type="bibr" rid="CIT416">Wan+19</xref>; <xref ref-type="bibr" rid="CIT250">Liu+19b</xref>]. This type of fully convolutional architecture is of common usage in image-space regression problems, due to its capability of efficiently learning patterns at different levels of abstraction thanks to its multi-scale design. Skip connections are added to enhance local details [<xref ref-type="bibr" rid="CIT344">RFB15</xref>; <xref ref-type="bibr" rid="CIT176">Iso+17</xref>]. The small number of parameters allows for faster training and inference, as well as reduced memory usage.</p>
<p>Note that, as opposed to previous work [<xref ref-type="bibr" rid="CIT067">DDB20</xref>] on material transfer, which is initialized from a pre-trained network, but inspired by single-sample image synthesis methods [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>; <xref ref-type="bibr" rid="CIT359">SDM19</xref>], we train one network per material and visual property. This strategy, although increasing training time, is key for the following reasons: First, it allows us to better understand the generalization capabilities obtained through our data augmentation policies. Second, it guarantees predictability of the trained model as every feature learned by the network is specific for each material and visual property pairs. Finally, it removes potential problems of a biased dataset as there is no cross-material <target target-type="page" id="pges_37"/>or cross-domain learning. If the input dataset does not contain enough variations of the material to represent the whole material, the model will fail in areas with unseen patterns. In those cases, having a pre-trained network may help as a material prior, as in [<xref ref-type="bibr" rid="CIT067">DDB20</xref>]. However, our method obtains comparable results to pre-trained methods, with a smaller computational footprint and with the additional flexibility of not needing an expensive dataset and large models. In practice, training a single network per material and attribute is not problematic as this process takes less than a minute.</p>
<p><bold>Loss Function</bold> Choosing the appropriate loss function for a learning framework highly depends on the problem. For example, some methods [<xref ref-type="bibr" rid="CIT176">Iso+17</xref>] combine a per-pixel <inline-formula><mml:math id="M66" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> metric with adversarial [<xref ref-type="bibr" rid="CIT116">Goo+14a</xref>] losses, as the latter allows for better semantic mappings in multi-modal learning scenarios, whilst the former allows for improved predictability. Texler <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>] further includes a perceptual loss [<xref ref-type="bibr" rid="CIT386">Tex+20a</xref>], while using a render loss is common in methods that estimate material parameters from photos [<xref ref-type="bibr" rid="CIT065">Des+18</xref>; <xref ref-type="bibr" rid="CIT067">DDB20</xref>]. In our method, we show that using a <inline-formula><mml:math id="M67" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> loss for training is enough for learning accurate and predictable mappings in our regression task, and binary cross-entropy loss for semantic segmentation, following standard practice in image segmentation [<xref ref-type="bibr" rid="CIT344">RFB15</xref>]. We found that the <inline-formula><mml:math id="M68" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> loss function yields overly-smooth outputs, and perceptual loss functions like <italic>LPIPS</italic> [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] are more prone to artifacts than pixel-wise <inline-formula><mml:math id="M69" display='block'><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mtext>norm</mml:mtext></mml:msub></mml:math></inline-formula> metrics. We include an ablation study of the impact of the loss function in the supplementary material of this chapter.</p>
</sec>
<sec id="c2-s4.s3">
<label>2.4.3</label>
<title>Implementation details</title>
<p>We use PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>] as the learning framework, Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>] for optimization, a learning rate of 0.002, and a batch size of 16. The training images are randomly augmented using uniform distributions by the following operations, in order: First, the photometric input is subject to the color augmentation policy. Then, both inputs and targets are randomly rotated by an angle in the [&#x2212;90, 90] range, randomly sheared by an angle in the [&#x2212;45, 45] range, and randomly rescaled in the [0.5, 2] range of scale factors. Then, patches of 128 &#x00D7; 128 pixels are randomly cropped during training to generate a large dataset of images.</p>
<p>All inputs are always standardized using their own mean and variance. Each <inline-formula><mml:math id="M70" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> is trained for 1000 iterations, which takes around 1 minute on a single Nvidia 1080Ti GPU. Due to the fully convolutional nature of our models and their reduced number of trainable parameters, the guidance images X used for evaluation can be of arbitrary dimensions. We used images of up to 5000 &#x00D7; 5000 pixels for which evaluation takes around 150ms. We refer the reader to <target target-type="page" id="pges_38"/>the supplementary material in this chapter for a comprehensive description and a diagram of this model, as well as further implementation details.</p>
</sec>
</sec>
<sec id="c2-s5">
<label>2.5</label>
<title>Dataset and Metrics</title>
<sec id="c2-s5.s1">
<label>2.5.1</label>
<title>Dual-Resolution Captured Data</title>
<p>Our method is agnostic to the capture setup [<xref ref-type="bibr" rid="CIT291">Nam+16</xref>; <xref ref-type="bibr" rid="CIT272">MWK17</xref>; <xref ref-type="bibr" rid="CIT007">Alc+19</xref>], and it may work for any kind of input data (e.g. BTFs [<xref ref-type="bibr" rid="CIT233">LM01</xref>; <xref ref-type="bibr" rid="CIT317">Rai+19</xref>]) as long as the photometric images are pixel-wise aligned with each other and to the visual attribute. This dataset could also be created synthetically by rendering the outputs of any SVBRDF estimation method [<xref ref-type="bibr" rid="CIT067">DDB20</xref>]. In <xref ref-type="sec" rid="c2-s7">Section 2.7</xref>, we show results of our method using these acquisition pipelines.</p>
<p>However, for the purpose of this evaluation, we choose to work with real data. The main reason is that the data obtained with material capture devices poses extra challenges that are difficult to reproduce with render engines. For example, material irregularities, complex optical behavior, or distortions and color shifts introduced by the optical system that might cause the models to produce inaccurate estimations.</p>
<p>We create a dataset containing images of the same material taken with two imaging camera systems: a high-resolution camera that allows us to take pictures of 0.7 &#x00D7; 0.9 cm, with a resolution of 367 &#x00D7; 490 pixels and a macroscopic camera which provides images of 11 &#x00D7; 11 cm, with a resolution of 4800 &#x00D7; 4800 pixels. In <target target-type="page" id="pges_39"/>terms of illumination, our setup has 27 different collimated light sources uniformly distributed across the hemisphere as well as diffuse illumination. We build a dataset of four different textile materials, whose complex optical behavior due to anisotropy, transmittance, directionality and microstructure [<xref ref-type="bibr" rid="CIT040">CLA19</xref>] turns them particularly challenging for synthesis and editing operations [<xref ref-type="bibr" rid="CIT191">Kam+16</xref>; <xref ref-type="bibr" rid="CIT317">Rai+19</xref>]. For training, we capture one image for each light source, making a total of 28 different images (<xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref> (a)). For the evaluation set, we take one <italic>guidance</italic> image with diffuse illumination and another <italic>guidance</italic> image with a directional light source (<xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref> (c)). The visual attributes (<xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref> (b)) are generated automatically using photometric stereo [<xref ref-type="bibr" rid="CIT174">Ike81</xref>] in the case of normal maps and manually by artists in the cases of colorizations and segmentation masks.</p>
</sec>
<sec id="c2-s5.s2">
<label>2.5.2</label>
<title>Attribute-specific Metrics for Evaluation</title>
<p>For evaluating our models, we choose domain-specific distance metrics different from those they were trained with to better understand their generalization capabilities [<xref ref-type="bibr" rid="CIT388">TOB15</xref>]. For normal maps, we compute the cosine distance between ground truth and estimated maps, as it accounts for the geometric space in which normals lie. In the case of image segmentation, we evaluate the results using the <italic>Jaccard similarity coefficient</italic> (IoU) [<xref ref-type="bibr" rid="CIT459">Zho+19</xref>], which is well-suited for sparse segmentation tasks. Finally, we use the state-of-the-art metric <italic>Learned Perceptual Image Patch Similarity</italic> (LPIPS) to evaluate the quality of the colorizations, as it has been shown to outperform <inline-formula><mml:math id="M71" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> norm for visual perception tasks [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>].</p>
</sec>
</sec>
<sec id="c2-s6">
<label>2.6</label>
<title>Evaluation</title>
<p>We evaluate our method in different settings. First, we assess the type and amount of photometric input data necessary for the model to generalize to different illumination conditions and image distortions. Then, we test our data augmentation strategy for affine transformations.</p>
<sec id="c2-s6.s1">
<label>2.6.1</label>
<title>Invariance to Input Illuminations and Distortions</title>
<p>In this set of experiments, we aim to evaluate if the model produces the same output after changing the training data and input guidance images. To this end, we measure both the impact of different illuminations and sizes of the photometric dataset, as well as variations of the input guidance images.</p>
<sec id="c2-s6.s1.1">
<title>Photometric Input</title>
<p>In this experiment, we study the type of photometric input that makes the method invariant to different illumination conditions of the guidance image. We <target target-type="page" id="pges_40"/>compare three models trained with different datasets: <italic>diffuseNet</italic>, which uses a single image illuminated with diffuse lighting; <italic>photometricNet</italic>, which takes as input 27 directional lights; and <italic>diphotoNet</italic>, which uses all the 28 sources. For evaluation, we use the macroscopic camera which captures a larger sample of the material at a lower resolution. We take two test guidance images with diffuse and directional lighting (<xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref> (c)). In this experiment, all images are aligned, therefore, we follow a limited policy of data augmentation from <xref ref-type="sec" rid="c2-s4.s1">Section 2.4.1</xref>, applying only color augmentation, rescaling, and crops, and leaving out rotations and shears for in-the-wild scenarios.</p>
<p><xref ref-type="fig" rid="fig-2.6">Figure 2.6</xref> shows qualitative results of our method for a selection of patches for the <italic>denim</italic>material. <xref ref-type="fig" rid="fig-2.6">Figure 2.6</xref> (a) shows that the best accuracy is in general obtained with <italic>diphotoNet</italic>, i.e. training with both directional and diffuse illuminations. This is reasonable, as the model is trained using the same illuminations used during test time. Conversely, <italic>diffuseNet</italic>, a network only trained using diffuse lighting is less capable of generalizing under any kind of illumination source. This suggests that following a photometric approach for training material synthesis models allows for better generalization capabilities. Finally, the results for <italic>photometricNet</italic>, which is trained only with directional lights shows similar accuracy as <italic>diphotoNet</italic> while proving generalization to unseen illuminations. For all the results shown in the chapter, we have chosen <italic>photometricNet</italic> as our model. This has an additional advantage from the usability perspective: it is relatively easy and cost-effective to build a capture setup (such as the flash light from a smartphone) or generate renders with directional lights, while it is considerably harder to recreate both types of illumination consistently.</p>
<fig id="fig-2.6">
<label>Figure 2.6:</label>
<caption><title>Qualitative results of our models under different datasets and data augmentation configurations, for different inputs of the <italic>denim</italic> material. On the left (a), we show the results of our networks trained using different dataset configurations, under a guidance image taken with diffuse lighting (<italic>p</italic><sub>0</sub>), and three crops of a guidance image illuminated with a directional light (<italic>p</italic><sub>1</sub>, <italic>p</italic><sub>2</sub>, <italic>p</italic><sub>3</sub>) not present in the training set. Please refer to <xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref> for the position of these crops on the larger guidance images. On the right (b), we show the results of our <italic>photometricNet</italic>, under different geometric distortions (rotations and shears) performed to its guidances images, taken under diffuse illumination.</title></caption>
<alt-text>Qualitative results of our models under different datasets and data augmentation config-urations, for different inputs of the denim material. On the left (a), we show the results of our networks trained using different dataset configurations, under a guidance image taken with diffuse lighting (p0), and three crops of a guidance image illuminated with a directional light (p1, p2, p3) not present in the training set. Please refer to Figure 2.5 for the position of these crops on the larger guidance images. On the right (b), we show the results of our photometricNet, under different geometric distortions (rotations and shears) performed to its guidances images, taken under diffuse illumination.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.6.jpg"/>
</fig>
</sec>
<sec id="c2-s6.s1.2">
<title>Dataset Size Influence</title>
<p>Our goal in this experiment is to understand how many images [<xref ref-type="bibr" rid="CIT295">NJR15</xref>] taken under directional lights are needed in order to obtain the desired invariance to illumination.</p>
<p>We train different models using reduced versions of our full 27-image dataset, and compare their results to those of the <italic>photometricNet</italic> trained on the full directional dataset. More precisely, we train networks using 1, 3, 9 and 18 directional lights for each material and application in our dataset following the same reduced data augmentation policy described in the previous experiment. Instead of randomly selecting light sources around the hemisphere, we perform a more sensible light source sampling. Specifically, for each reduced dataset, there is at least one light that is as close as possible to the normal of the surface on which the material lies, thus giving more importance to frontal angles. We extend these selected lights to 3 for the reduced datasets with more than three lights. The rest of lighting sources are uniformly sampled around the hemisphere, up to a zenith angle of 70 degrees. The supplementary material contains a diagram of the distribution of lights in the hemisphere. Results are shown in <xref ref-type="fig" rid="fig-2.7">Figure 2.7</xref>, where we see <target target-type="page" id="pges_41"/>that adding more lights to the dataset monotonically increases the generalization of every model. However, adding lights consistently shows diminishing returns, which suggests that a capture setup with around nine lights might be enough for relatively accurate estimations. Similar findings are reported in multi-image SVBRDF estimation methods [<xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT066">Des+19</xref>].</p>
<fig id="fig-2.7">
<label>Figure 2.7:</label>
<caption><title>Error of the reduced <italic>photometricNets</italic> on our ground truth guidance image (taken with diffuse illumination, not present in the training dataset) for each material in the dataset and the three visual attributes and corresponding error metrics, which are close to the normal of the surface in which the material lies. The width of the lines indicates the standard variation across 5 repeated experiments, where different light directions were randomly chosen to form the training datasets. Please refer to the supplementary material for the position of each light source.</title></caption>
<alt-text>Error of the reduced photometricNets on our ground truth guidance image (taken with diffuse illumination, not present in the training dataset) for each material in the dataset and the three visual attributes and corresponding error metrics, which are close to the normal of the surface in which the material lies. The width of the lines indicates the standard variation across 5 repeated experiments, where different light directions were randomly chosen to form the training datasets. Please refer to the supplementary material for the position of each light source.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.7.jpg"/>
</fig>
</sec>
<sec id="c2-s6.s1.3">
<title>Image Degradations</title>
<p>As discussed in <xref ref-type="sec" rid="c2-s5.s1">Section 2.5.1</xref>, real captured images may be subject to distortions, shifts and noise introduced by the optical capture system. To fully understand the robustness of our models with respect to these types of imperfections, we synthetically modify the saturation, contrast and noise present in the X guidance images. As shown in <xref ref-type="fig" rid="fig-2.8">Figure 2.8</xref>, <inline-formula><mml:math id="M72" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> is robust to these types of degradations, even when a fair amount of details are lost in X. These images were not part of the training dataset.</p>
<fig id="fig-2.8">
<label>Figure 2.8:</label>
<caption><title>Output of our model under different distorted inputs: a change of saturation, contrast, and Gaussian noise (&#x03C3;<sup>2</sup> = 255). The number at the bottom is the cosine distance with respect to the original estimation. As we can see, the output is consistent with a very small error in all cases.</title></caption>
<alt-text>Output of our model under different distorted inputs: a change of saturation, contrast, and Gaussian noise (&#x03C3;2 = 255). The number at the bottom is the cosine distance with respect to the original estimation. As we can see, the output is consistent with a very small error in all cases.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.8.jpg"/>
</fig>
</sec>
</sec>
<sec id="c2-s6.s2">
<label>2.6.2</label>
<title>Equivariance to Affine Transforms</title>
<p>Our goal in this experiment is to analyze the equivariance of our model to affine distortions of the guidance images. A function <italic>f</italic> is said to be <italic>equivariant</italic> with respect to a transformation <inline-formula><mml:math id="M73" display='block'><mml:mi>T</mml:mi><mml:mtext> if </mml:mtext><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Inspired by [<xref ref-type="bibr" rid="CIT027">Ben+20</xref>], we measure the equivariance of our model <inline-formula><mml:math id="M74" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> with respect to different affine transforms <inline-formula><mml:math id="M77" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> performed to an input guidance image X, by computing their difference using the corresponding metric defined in <xref ref-type="sec" rid="c2-s5.s2">Section 2.5.2</xref> (<inline-formula><mml:math id="M78" display='block'><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="script">T</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="script">T</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. To achieve this, we extend the augmentation policy in the previous experiments by adding random shears and rotations to the training process. We train the <italic>photometricNet</italic> with and without random shears and rotations for every visual attribute in our dataset and on the <italic>denim</italic> material. In the case of the normals, as described in <target target-type="page" id="pges_42"/><xref ref-type="sec" rid="c2-s4.s1">Section 2.4.1</xref>, we also perform the shears and rotations in the geometric space in which normals lie.</p>
<p>We measure their robustness with respect to three different transformations <inline-formula><mml:math id="M79" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> : rescalings, rotations and shears. As before, we use a guidance image taken under diffuse lighting as input for these experiments and attribute-specific distance metrics. <xref ref-type="fig" rid="fig-2.6">Figure 2.6</xref> (b) shows that not adding those transforms generates visual artifacts for rotations and shears. Interestingly, without affine augmentations, the models hallucinates vertical yarns, as it is the only type of data that it has seen as input. <xref ref-type="fig" rid="fig-2.9">Figure 2.9</xref> shows quantitative metrics for the range of transformations <inline-formula><mml:math id="M80" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> in which we evaluated the models. Augmenting the training dataset with both shears and rotations generally provides the best results. Furthermore, this enhanced data augmentation policy improves the robustness of the models with respect to the scale of their inputs. Notably, applying random rotations during training appears to have a larger impact than random shears on the robustness of the models. This finding suggests that applying blind policies of data augmentation [<xref ref-type="bibr" rid="CIT012">ASE17</xref>; <xref ref-type="bibr" rid="CIT363">SK19</xref>; <xref ref-type="bibr" rid="CIT350">San+19</xref>; <xref ref-type="bibr" rid="CIT047">Che+20b</xref>] may not be an optimal strategy for some applications like ours, as the networks may not learn relevant features, or overfit to noise present in the training dataset. Finally, it is worth noting that every network has the same number of parameters, and is trained for the same number of iterations, which shows that the generalization capabilities of the models can be increased at no extra parameter cost.</p>
<fig id="fig-2.9">
<label><target target-type="page" id="pges_43"/>Figure 2.9:</label>
<caption><title>Quantitative evaluation of the robustness of our networks with respect to different transformations <inline-formula><mml:math id="M81" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> (rescaling, rotation and shearing) of networks trained under different data augmentation policies for the <italic>denim</italic> material. From top to bottom, we show results on: recoloring, segmentation, and normals estimation. Lower is better for each metric</title></caption>
<alt-text>Quantitative evaluation of the robustness of our networks with respect to different transformations T (rescaling, rotation and shearing) of networks trained under different data augmentation policies for the denim material. From top to bottom, we show results on: recoloring, segmentation, and normals estimation. Lower is better for each metric</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.9.jpg"/>
</fig>
</sec>
</sec>
<sec id="c2-s7">
<label>2.7</label>
<title>Results and Comparisons</title>
<p>In this section, we first compare our method with related approaches on image stylization and large scale SVBRDF material transfer. Then, we show the capabilities of our method to generalize to similar materials to those in their training set and present its limitations.</p>
<sec id="c2-s7.s1">
<label>2.7.1</label>
<title>Interactive Stylizations</title>
<p>In the first category, the method of Texler <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT386">Tex+20a</xref>] allows artists to interactively edit a few keyframes of a video and propagate that edition to the rest of the video. As ours, they formulate this transfer problem by training an encoder-decoder network using patch-based learning. But, in contrast to our work, they do not perform any data augmentation policy outside of random cropping. For a fair comparison with such method, we compare two setups: 1) using a single image as training input, and 2) using the photometric dataset. Results are shown in <xref ref-type="fig" rid="fig-2.2">Figures 2.2</xref> and <xref ref-type="fig" rid="fig-2.10">2.10</xref>. As Texler&#x2019;s method is not scale invariant, in both cases the training data provided for their model has the same scale as the images used for testing. <target target-type="page" id="pges_44"/>Our models have been trained with the full policy of data augmentation. In the first setup (<xref ref-type="fig" rid="fig-2.10">Figure 2.10</xref>), we use the diffuse illumination and hence compare their output with our <italic>diffuseNet</italic> output. As shown, none of the methods provide high quality results but our model manages to provide closer estimations. In the second setup (<xref ref-type="fig" rid="fig-2.2">Figure 2.2</xref>), we train their model with the photometric input. The best results are obtained with our method. Even though extending the input data using photometric cues has an impact on the quality of Texler&#x2019;s results, the lack of a data augmentation policy makes the transfer fuzzier and noisier. Further, their combination of style, adversarial and pixel-wise losses fails to yield predictable mappings. These results confirm the importance of a comprehensive data-augmentation policy, such as the one we propose, when using neural networks for image processing tasks of this kind. In addition, our model is trained in less time with a smaller computational footprint (1 minute vs 5 minutes).</p>
<fig id="fig-2.10">
<label>Figure 2.10:</label>
<caption><title>Comparison of our method with Texler <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>] using a single diffuse image of the <italic>denim</italic> material as training data. The task is to transfer the two attributes shown (<italic>Stylization</italic> and <italic>Segmentation</italic>) to two different guidance images. Even using a single image instead of a photometric dataset, we achieve higher quality mappings at a lower cost.</title></caption>
<alt-text>Comparison of our method with Texler et al. [Tex+20b] using a single diffuse image of the denim material as training data. The task is to transfer the two attributes shown (Stylization and Segmentation) to two different guidance images. Even using a single image instead of a photometric dataset, we achieve higher quality mappings at a lower cost.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.10.jpg"/>
</fig>
<p><target target-type="page" id="pges_45"/>Another way of formulating this visual attribute transfer problem is through <italic>image analogies</italic>. Using our single diffuse image for input, we compare our approach with two methods as shown in <xref ref-type="fig" rid="fig-2.11">Figure 2.11</xref>. First, the work of Liao <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>], which uses deep latent spaces as image descriptors; and the method of Benaim <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>] that trains single-image generative models to find bijective mappings between the <italic>structure</italic> of one image and the <italic>style</italic> of another. Our method qualitatively outperforms these methods with a fraction of the computational cost: 1 minute in our case, 10 minutes in [<xref ref-type="bibr" rid="CIT150">Her+01</xref>], 40 minutes in [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>] and 10 hours in [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>]. Once trained, our models can be used to evaluate any guidance image in real time for materials with similar microstructure. In contrast, image analogies methods require expensive optimizations for each guidance image. We refer the reader to the supplementary material for more comparisons with these methods.</p>
<fig id="fig-2.11">
<label><target target-type="page" id="pges_46"/>Figure 2.11:</label>
<caption><title>Comparison of image analogies approaches for input shown on the left, of the <italic>denim</italic> material. From left to right: Deep Image Analogies [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>] and Structural Analogies [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>]. Our results shows more accurate and predictable mappings, at less computational cost than the alternatives.</title></caption>
<alt-text>Comparison of image analogies approaches for input shown on the left, of the denim material. From left to right: Deep Image Analogies [Lia+17] and Structural Analo-gies [Ben+21]. Our results shows more accurate and predictable mappings, at less computational cost than the alternatives.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.11.jpg"/>
</fig>
<p><xref ref-type="fig" rid="fig-2.12">Figure 2.12</xref> shows additional results of material stylizations. In these examples, we used the method of Gatys <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT108">GEB15b</xref>] to stylize a small patch of the material. Then, we trained a model using photometricNet and diffuseNet. As the guidance image we used a bigger image with diffuse illumination. Compared with na&#x00EF;ve style transfer applied to the whole image, our approach provides detailed stylizations where the microstructure of the material is preserved. We further see that diffuseNet provides noisier results than photometricNet, probably due to the fact that the photometric cues help to preserve the local shading variations.</p>
<fig id="fig-2.12">
<label>Figure 2.12:</label>
<caption><title>Our method allows for interactive material-aware visual attribute transfers. Using an off-the-shelf style transfer algorithm [<xref ref-type="bibr" rid="CIT108">GEB15b</xref>], we can transfer the style of one image <inline-formula><mml:math id="M82" display='block'><mml:msub><mml:mi>I</mml:mi><mml:mtext>style</mml:mtext></mml:msub></mml:math></inline-formula> to the content of another <inline-formula><mml:math id="M83" display='block'><mml:msub><mml:mi>I</mml:mi><mml:mtext>content</mml:mtext></mml:msub></mml:math></inline-formula>, obtaining a visual attribute <italic>w</italic>. Training a <inline-formula><mml:math id="M84" display='block'><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>&#x03C9;</mml:mi></mml:msub></mml:math></inline-formula> to learn this relationship, we can find predictable style mappings, that we can transfer to guidance images X, obtaining style transfers <inline-formula><mml:math id="M85" display='block'><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>&#x03C9;</mml:mi></mml:msub></mml:math></inline-formula>(X). Learning this transfer is inexpensive and allows for interactive editions. Performing this transfer directly to the guidance image generates artifacts and not-predictable mappings.</title></caption>
<alt-text>Our method allows for interactive material-aware visual attribute transfers. Using an off-the-shelf style transfer algorithm [GEB15b], we can transfer the style of one image Istyle to the content of another Icontent, obtaining a visual attribute w. Training a M&#x03C9; to learn this relationship, we can find predictable style mappings, that we can transfer to guidance images X, obtaining style transfers M&#x03C9;(X). Learning this transfer is inexpensive and allows for interactive editions. Performing this transfer directly to the guidance image generates artifacts and not-predictable mappings.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.12.jpg"/>
</fig>
</sec>
<sec id="c2-s7.s2">
<label>2.7.2</label>
<title>Creation of Large Scale Digital Material Assets</title>
<p>Our method can be used to propagate SVBRDFs estimated locally in a small area of the material to larger samples. <xref ref-type="fig" rid="fig-2.13">Figure 2.13</xref> illustrates for a diverse set of materials that we can propagate albedo and normals estimated at high resolution in a small area of 0.7 x 0.9 cm to guidance images of 13 x 13 cm taken with a smartphone. Even though we train the albedo and normals models separately, which does not guarantee pixel-wise coherence between the estimated maps, the rendered images show realistic-looking materials, even with a diffuse material model. As opposed to the method of Deschaintre <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT067">DDB20</xref>], which is trained to output directly Cook-Torrance [<xref ref-type="bibr" rid="CIT056">CT82</xref>] material layers, our method is agnostic to the parameters of the SVBRDF.</p>
<fig id="fig-2.13">
<label><target target-type="page" id="pges_47"/>Figure 2.13:</label>
<caption><title>Results of our framework for material capture using a smartphone. Training two models, <inline-formula><mml:math id="M86" display='block'><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M87" display='block'><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:math></inline-formula> with a photometric dataset, we obtain respectively albedo and normals from guidance images taken under uncontrolled conditions, which can be used by render engines. Here we have used Arnold [<xref ref-type="bibr" rid="CIT114">Geo+18</xref>] and a diffuse material model.</title></caption>
<alt-text>Results of our framework for material capture using a smartphone. Training two models, Ma and Mn with a photometric dataset, we obtain respectively albedo and normals from guidance images taken under uncontrolled conditions, which can be used by render engines. Here we have used Arnold [Geo+18] and a diffuse material model.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.13.jpg"/>
</fig>
<p>We assess the capabilities of our method on this setting using the same SVBRDF propagation scenario as proposed in [<xref ref-type="bibr" rid="CIT067">DDB20</xref>]. Using a small crop of a synthetic SVBRDF as input, we render 27 images using the same directional light position as we used in our real dataset, and train a photometricNet to estimate the surface normals from each of those renders. We then evaluate this model using a larger area of the material, illuminated under an unknown lighting position. In <xref ref-type="table" rid="c2-tab1">Table 2.1</xref>, we show a quantitative comparison <target target-type="page" id="pges_48"/>with [<xref ref-type="bibr" rid="CIT067">DDB20</xref>], under different image quality metrics. As shown, our method achieves better scores on pixel-wise metrics, whilst [<xref ref-type="bibr" rid="CIT067">DDB20</xref>] achieves better deep perceptual scores, as in the LPIPS metric [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>]. This might be related to the design of our loss function: we directly minimize pixel-wise differences, while [<xref ref-type="bibr" rid="CIT067">DDB20</xref>] is optimized using a render-aware loss. Qualitatively, as shown in <xref ref-type="fig" rid="fig-2.14">Figure 2.14</xref>, our method obtains comparable quality mappings, with fewer artifacts. Despite capturing SVBRDF or arbitrary materials is not the goal of our method, we include in the supplementary comparisons with a direct SVBRDF acquisition method [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>] using our own real dataset.</p>
<table-wrap id="c2-tab1">
<label>Table 2.1:</label>
<caption><title>Quantitative comparison with [<xref ref-type="bibr" rid="CIT067">DDB20</xref>], on the studied materials and different performance metrics. As shown, our method provides better pixel-wise accuracy than [<xref ref-type="bibr" rid="CIT067">DDB20</xref>], while their method obtains better perceptual scores.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left" rowspan="2"><p><bold>Material ID</bold></p></th>
<th valign="top" align="left" colspan="2"><p><bold>SSIM</bold> [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>]&#x2191;</p></th>
<th valign="top" align="center" colspan="2"><p><bold>PSNR</bold> &#x2191;</p></th>
<th valign="top" align="center" colspan="2"><p><bold>MSE</bold> &#x2193;</p></th>
<th valign="top" align="left" colspan="2"><p><bold>LPIPS</bold> [<xref ref-type="bibr" rid="CIT450">Zha+ 18</xref>] &#x2193;</p></th>
</tr>
<tr>
<th valign="top" align="left"><p>[<xref ref-type="bibr" rid="CIT067">DDB20</xref>]</p></th>
<th valign="top" align="left"><p><bold>Ours</bold></p></th>
<th valign="top" align="left"><p>[<xref ref-type="bibr" rid="CIT067">DDB20</xref>]</p></th>
<th valign="top" align="left"><p><bold>Ours</bold></p></th>
<th valign="top" align="left"><p>[<xref ref-type="bibr" rid="CIT067">DDB20</xref>]</p></th>
<th valign="top" align="left"><p><bold>Ours</bold></p></th>
<th valign="top" align="left"><p>[<xref ref-type="bibr" rid="CIT067">DDB20</xref>]</p></th>
<th valign="top" align="left"><p><bold>Ours</bold></p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p>560</p></td>
<td valign="top" align="left"><p>0,905</p></td>
<td valign="top" align="left"><p>0,925</p></td>
<td valign="top" align="left"><p>31,190</p></td>
<td valign="top" align="left"><p>32,310</p></td>
<td valign="top" align="left"><p>0,002</p></td>
<td valign="top" align="left"><p>0,002</p></td>
<td valign="top" align="left"><p>0,244</p></td>
<td valign="top" align="left"><p>0,221</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>1581</p></td>
<td valign="top" align="left"><p>0,645</p></td>
<td valign="top" align="left"><p>0,675</p></td>
<td valign="top" align="left"><p>29,740</p></td>
<td valign="top" align="left"><p>29,920</p></td>
<td valign="top" align="left"><p>0,006</p></td>
<td valign="top" align="left"><p>0,005</p></td>
<td valign="top" align="left"><p>0,398</p></td>
<td valign="top" align="left"><p>0,446</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>1684</p></td>
<td valign="top" align="left"><p>0,654</p></td>
<td valign="top" align="left"><p>0,746</p></td>
<td valign="top" align="left"><p>29,711</p></td>
<td valign="top" align="left"><p>30,410</p></td>
<td valign="top" align="left"><p>0,007</p></td>
<td valign="top" align="left"><p>0,003</p></td>
<td valign="top" align="left"><p>0,196</p></td>
<td valign="top" align="left"><p>0,251</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>2111</p></td>
<td valign="top" align="left"><p>0,770</p></td>
<td valign="top" align="left"><p>0,783</p></td>
<td valign="top" align="left"><p>32,250</p></td>
<td valign="top" align="left"><p>32,881</p></td>
<td valign="top" align="left"><p>0,002</p></td>
<td valign="top" align="left"><p>0,002</p></td>
<td valign="top" align="left"><p>0,397</p></td>
<td valign="top" align="left"><p>0,383</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Average</p></td>
<td valign="top" align="left"><p>0,744</p></td>
<td valign="top" align="left"><p>0,782</p></td>
<td valign="top" align="left"><p>30,723</p></td>
<td valign="top" align="left"><p>31,380</p></td>
<td valign="top" align="left"><p>0,004</p></td>
<td valign="top" align="left"><p>0,003</p></td>
<td valign="top" align="left"><p>0,309</p></td>
<td valign="top" align="left"><p>0,325</p></td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-2.14">
<label>Figure 2.14:</label>
<caption><title>Comparison of our method with the Guided Fine Tuning, by Deschaintre <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT067">DDB20</xref>]. Following their algorithm, we render the <italic>Input SVBRDFs</italic>, and train a photometricNet on those renders. As shown, our method can achieve higher quality normal maps, with fewer artifacts. Input SVBRDFs, X, ground truth and results from Guided Fine Tuning were obtained directly from [<xref ref-type="bibr" rid="CIT067">DDB20</xref>].</title></caption>
<alt-text>Comparison of our method with the Guided Fine Tuning, by Deschaintre et al. [DDB20]. Following their algorithm, we render the Input SVBRDFs, and train a photometricNet on those renders. As shown, our method can achieve higher quality normal maps, with fewer artifacts. Input SVBRDFs, X, ground truth and results from Guided Fine Tuning were obtained directly from [DDB20].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.14.jpg"/>
</fig>
<p>A similar setting can be used to propagate attributes leveraging BTF measurements as training data. <xref ref-type="fig" rid="fig-2.15">Figure 2.15</xref> shows an example using a BTF from [<xref ref-type="bibr" rid="CIT422">WGK14</xref>]. Using Photometric Stereo [<xref ref-type="bibr" rid="CIT174">Ike81</xref>], we compute their surface normals and train a photometricNet on a central crop of the BTF, using the captures at a camera position of (&#x03D5; = 0&#x00B0;, &#x03B8; = 0&#x00B0;). We then evaluate this model using the full <target target-type="page" id="pges_49"/>material surface, and a novel camera position, of (&#x03D5; = 15&#x00B0;, &#x03B8; = 11&#x00B0;). As shown, our model is capable of working with captured BTF data.</p>
<fig id="fig-2.15">
<label>Figure 2.15:</label>
<caption><title>Our method can work with real BTF captured data. Training a photometricNet on a crop of the material (summarized in <italic>Training Dataset</italic>), to output surface normals, we can estimate surface normals on larger areas of the material, even under novel viewing positions.</title></caption>
<alt-text>Our method can work with real BTF captured data. Training a photometricNet on a crop of the material (summarized in Training Dataset), to output surface normals, we can estimate surface normals on larger areas of the material, even under novel viewing positions.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.15.jpg"/>
</fig>
</sec>
<sec id="c2-s7.s3">
<label>2.7.3</label>
<title>Generalization to Similar Materials</title>
<p>In previous experiments, we have shown the performance of our models when the guidance image corresponds to the material used for training. In this experiment, we show that our models, despite being trained only on a set of captures of a single material, generalize to materials of the same category. <xref ref-type="fig" rid="fig-2.16">Figure 2.16</xref> shows some examples for models trained using our dataset (<xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref>), taking input guidance images of different albedos and scales. The transfer works thanks to the network design and data augmentation strategy that is designed to use microstructure details as guidance. Our method could thus be used to transfer visual attributes for a diverse set of materials by simply training one model using a single but representative material of each category.</p>
<fig id="fig-2.16">
<label>Figure 2.16:</label>
<caption><title>Generalization capabilities of our method when evaluated on materials similar to those in their training dataset. On the top row, we show the outputs of the model trained on the knit in our dataset (<xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref>), and evaluated on a different guidance images. We also show examples on linen and denim fabrics, with different conditions of saturation, blur, scale, and illumination. Even in very challenging cases, where the structure of the material is barely visible, as in the overly-saturated red linen or the noisy blue knit fabrics, our photometricNets can yield plausible results. The insets represent the training dataset by each model, as represented in <xref ref-type="fig" rid="fig-2.5">Figure 2.5</xref>.</title></caption>
<alt-text>Generalization capabilities of our method when evaluated on materials similar to those in their training dataset. On the top row, we show the outputs of the model trained on the knit in our dataset (Figure 2.5), and evaluated on a different guidance images. We also show examples on linen and denim fabrics, with different conditions of saturation, blur, scale, and illumination. Even in very challenging cases, where the structure of the material is barely visible, as in the overly-saturated red linen or the noisy blue knit fabrics, our photometricNets can yield plausible results. The insets represent the training dataset by each model, as represented in Figure 2.5.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.16.jpg"/>
</fig>
</sec>
<sec id="c2-s7.s4">
<label>2.7.4</label>
<title>Limitations</title>
<p>Our models are not guaranteed to provide high-quality results outside the range of input data and data augmentation policies we train them on. This limitation is common to all learning-based approaches. It is unlikely that the framework is capable of generalizing to resolutions higher to those of the training data; or down-sampled images in which the texture details are not recognizable. Furthermore, the type of transformations we apply to the training data may not represent all the possible geometric variations, or non-linear warpings that materials are subject to in the real world. For materials which exhibit a strong variation in their microstructure, which cannot be fully captured using a single photometric dataset, our patch-based approach will likely fail to generalize to the full heterogeneity. <xref ref-type="fig" rid="fig-2.17">Figure 2.17</xref> shows two examples of failure cases for which the test data is not included in the training data and where the affine and illumination transformations are outside the suitable range.</p>
<fig id="fig-2.17">
<label>Figure 2.17:</label>
<caption><title>Failure cases of our method. As shown in the first row, if the input dataset does not represent the heterogeneity present in the guidance X, the model fails to yield compelling results on unseen structures of the material, as shown on the green box. On the second row, we show a guidance image of the <italic>denim</italic> material X which exhibits strong geometric and illumination variations, outside of the range in which we train <inline-formula><mml:math id="M88" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> with. As such, the model shows a poor performance on the segmentation task.</title></caption>
<alt-text>Failure cases of our method. As shown in the first row, if the input dataset does not represent the heterogeneity present in the guidance X, the model fails to yield compelling results on unseen structures of the material, as shown on the green box. On the second row, we show a guidance image of the denim material X which exhibits strong geometric and illumination variations, outside of the range in which we train M with. As such, the model shows a poor performance on the segmentation task.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.17.jpg"/>
</fig>
</sec>
</sec>
<sec id="c2-s8">
<label>2.8</label>
<title>Conclusions and Future Work</title>
<p>In this chapter, we have proposed a neural visual attribute transfer framework capable of transferring, for a given material, many types of visual property maps to images of unseen patches of the same -or similar- material taken under different illumination, capture setup, and affine distortion. To our knowledge, the proposed framework is the first method capable of leveraging the optical behavior of the material to this purpose by being trained using a photometric approach. Such an approach, besides the illumination-invariance we have shown, helps <target target-type="page" id="pges_50"/>the neural network learn better mappings between visual domains, finding a physically-based representation of the material. Further, we have presented a comprehensive policy of data augmentation which outperforms previous work on visual attribute transfer given a single image of the material.</p>
<p>We have shown that our method can be used to transfer any kind of visual attribute estimated locally to larger material samples. Further, we have demonstrated that our models, although trained on a single material, generalize to materials of the same category. We think our findings will inspire future work showing that smart training strategies might alleviate the need for massive datasets. Our method could be extended in several ways. The need for obtaining high-resolution captures taken under <target target-type="page" id="pges_51"/>different illuminations may be reduced by generating rendering images through recent advances in inverse material acquisition [<xref ref-type="bibr" rid="CIT076">Don19</xref>]. Further training with multiple patches may help to cope with material heterogeneity [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>]. Similarly, our findings suggest that extending the data augmentation policy to include 3<italic>D</italic> deformations will likely improve the accuracy. Beyond the generation of large scale digital assets for rendering, our method may have potential in other visual computing applications that require a low level understanding of the properties of the materials in real scenes. For example, the yarn segmentation application shown in the chapter might be suitable as input to shape from texture applications. Specific visual attributes might be useful to identify or highlight defects for image forensics problems, or to enhance different features in real-time or AR applications.</p>
</sec>
<sec id="c2-sA">
<label><target target-type="page" id="pges_52"/>2.A</label>
<title>Additional Implementation Details</title>
<sec id="c2-sA.1">
<title>Network design</title>
<p>Our model, summarized in <xref ref-type="fig" rid="fig-2.18">Figure 2.18</xref>, is a standard U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] encoder-decoder network with 4 blocks of layers and skip connections. U-Nets are widely used in pixel-wise regression and image-to-image translation problems, due to their efficient design, capable of exploiting texture and semantic patterns at different levels of abstraction, due to their multi-scale design. Skip connections are known to significantly enhance the quality of deep image-to-image translation models [<xref ref-type="bibr" rid="CIT176">Iso+17</xref>]. Each block is comprised of a tuple of convolutional layers with kernel sizes of 3 &#x00D7; 3, followed by a batch normalization operation [<xref ref-type="bibr" rid="CIT175">IS15</xref>] and a ReLU [<xref ref-type="bibr" rid="CIT290">NH10</xref>] non-linearity. The number of convolutional filters in each layer is shown in <xref ref-type="fig" rid="fig-2.18">Figure 2.18</xref>. The last layer of the network uses 1 &#x00D7; 1 convolutional filters to transform the feature vectors into the desired number of classes (e.g 3 output channels for RGB regression tasks, as in normal map estimation or recoloring; or 1 output channel for the semantic segmentation problem). The total number of trainable parameters depends on the number of property maps that the network should jointly learn, but are around 483000 for all the problems we showcased in this work. Trainable weights are initialized randomly, by sampling a <inline-formula><mml:math id="M89" display='block'><mml:mi mathvariant="script">N</mml:mi><mml:mo>(</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.02</mml:mn><mml:mo>)</mml:mo></mml:math></inline-formula>. The output of each block in the encoder part of the network is max-pooled to halve its spatial resolution. Upsampling in the decoder side of the network is done through up-convolutions. For image segmentation problems, we add a Sigmoid non-linearity to the output of the network. Due to the reduced number of trainable parameters, models take only 1.94<italic>Mbs</italic>, allowing for fast training and evaluation. On <xref ref-type="table" rid="c2-tab2">Table 2.2</xref>, we show a quantitative study on the influence of the depth of the model in its accuracy, using the 1684 material from [<xref ref-type="bibr" rid="CIT065">Des+18</xref>] as the training dataset. A visualization of these results is included in <xref ref-type="fig" rid="fig-2.19">Figure 2.19</xref>, showing that deeper models yield overly-smooth results for this few-shot learning problem.</p>
<fig id="fig-2.18">
<label><target target-type="page" id="pges_53"/>Figure 2.18:</label>
<caption><title>An overview of our deep learning architecture <inline-formula><mml:math id="M90" display='block'><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>&#x03C9;</mml:mi></mml:msub></mml:math></inline-formula>, which is a shallow U-net net-work [<xref ref-type="bibr" rid="CIT344">RFB15</xref>]. Below each block of layers, we show the number of trainable convolutional kernels, as well as the spatial resolution spanned by each block, with respect to the input resolution <italic>D</italic>. We use the normal map estimation problem for a denim fabric for visualization purposes.</title></caption>
<alt-text>An overview of our deep learning architecture M&#x03C9;, which is a shallow U-net net-work [RFB15]. Below each block of layers, we show the number of trainable convolutional kernels, as well as the spatial resolution spanned by each block, with respect to the input resolution D. We use the normal map estimation problem for a denim fabric for visualization purposes.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.18.jpg"/>
</fig>
<table-wrap id="c2-tab2">
<label>Table 2.2:</label>
<caption><title>Quantitative comparison between different number of blocks of layers in our model. Every model we use throughout this work has 4 blocks of layers, as it provides the best trade-off between computational cost and precision. Deeper models fail to generalize properly in this few-shot learning scenario. We use a color code to highlight <bold>best</bold> and <bold>worst</bold> cases</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left"><p>Depth</p></th>
<th valign="top" align="left"><p>SSIM [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>] &#x2191;</p></th>
<th valign="top" align="left"><p>PSNR &#x2191;</p></th>
<th valign="top" align="left"><p>MSE &#x2193;</p></th>
<th valign="top" align="left"><p>LPIPS [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] &#x2193;</p></th>
<th valign="top" align="left"><p>Size (Mbs) &#x2193;</p></th>
<th valign="top" align="left"><p>Train Time (s) &#x2193;</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p>3</p></td>
<td valign="top" align="left"><p><bold>0.71</bold></p></td>
<td valign="top" align="left"><p><bold>29.83</bold></p></td>
<td valign="top" align="left"><p><bold>0.0052</bold></p></td>
<td valign="top" align="left"><p><bold>0.37</bold></p></td>
<td valign="top" align="left"><p><bold>0.49</bold></p></td>
<td valign="top" align="left"><p><bold>51</bold></p></td>
</tr>
<tr>
<td valign="top" align="left"><p><bold>4</bold></p></td>
<td valign="top" align="left"><p><bold>0.74</bold></p></td>
<td valign="top" align="left"><p>30.41</p></td>
<td valign="top" align="left"><p><bold>0.0031</bold></p></td>
<td valign="top" align="left"><p><bold>0.25</bold></p></td>
<td valign="top" align="left"><p>1.94</p></td>
<td valign="top" align="left"><p>58</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>5</p></td>
<td valign="top" align="left"><p>0.72</p></td>
<td valign="top" align="left"><p>30.37</p></td>
<td valign="top" align="left"><p>0.0043</p></td>
<td valign="top" align="left"><p>0.33</p></td>
<td valign="top" align="left"><p>7.63</p></td>
<td valign="top" align="left"><p>164</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>6</p></td>
<td valign="top" align="left"><p>0.72</p></td>
<td valign="top" align="left"><p><bold>30.87</bold></p></td>
<td valign="top" align="left"><p>0.0041</p></td>
<td valign="top" align="left"><p>0.31</p></td>
<td valign="top" align="left"><p><bold>30.49</bold></p></td>
<td valign="top" align="left"><p><bold>283</bold></p></td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-2.19">
<label>Figure 2.19:</label>
<caption><title>Impact of the depth of the model on its accuracy, using a synthetic dataset. As shown, our baseline outperforms deeper and shallower configurations.</title></caption>
<alt-text>Impact of the depth of the model on its accuracy, using a synthetic dataset. As shown, our baseline outperforms deeper and shallower configurations.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.19.jpg"/>
</fig>
</sec>
<sec id="c2-sA.2">
<title>Training details</title>
<p>We use PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>] as the learning framework, Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>] for optimization, a learning rate of <inline-formula><mml:math id="M91" display='block'><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.002</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.9</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.999</mml:mn><mml:mo>,</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>8</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and a batch size of 16. The training images are randomly augmented using a uniform distribution by the following operations, in order: randomly rotated by an angle in the [&#x2212;90, 90] range, randomly sheared by an angle in the [&#x2212;45, 45] range, and randomly rescaled in the [0.5, 2] range of scale factors. Their color channels are randomly reordered and patches of 128 &#x00D7; 128 pixels are randomly cropped during training, to generate a large dataset of images. The rescaling factors and rotation angles are randomly chosen for each image in each batch using a uniform distribution on the mentioned ranges. All these data augmentation <target target-type="page" id="pges_54"/>operations are also applied to the target visual attributes, except for the color channel reordering, which is only applied to the input images. Images and visual attributes are always standardized using their own mean and variance. All the data augmentation operations are performed natively in GPU, which allows for faster training. To further reduce the computational cost and training times of our models, we train our models using automatic mixed precision training [<xref ref-type="bibr" rid="CIT274">Mic+18</xref>], which reduces the memory footprint of the models in GPU, while allowing for faster computation of certain operations, such as convolutions.</p>
<p>Each <inline-formula><mml:math id="M92" display='block'><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>&#x03C9;</mml:mi></mml:msub></mml:math></inline-formula> is trained independently for each material and application (no cross-domain training or fine-tuning) for 1000 iterations, which takes around 1 minute on a single 1080Ti GPU. Due to the fully convolutional nature of our models, and their reduced number of trainable parameters, the guidance images X used for evaluation can be of arbitrary resolutions. We evaluate the models using half-precision, which allows us to utilize guidance images of up to 5000 &#x00D7; 5000 pixels, for which evaluation takes around 150ms in GPU. Thanks to the reduced number of parameters, an efficient use of GPU-native data augmentation techniques and leveraging half precision training and evaluation, we are able to train and evaluate our photometricNets in less than a minute of total computational time.</p>
</sec>
<sec id="c2-sA.3">
<title>Ablation Study of the Loss Function</title>
<p>Finally, in <xref ref-type="fig" rid="fig-2.20">Figure 2.20</xref>, we show a comparison of the results of our models when trained on different loss functions. As shown, even though the differences are sometimes subtle, the <inline-formula><mml:math id="M93" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> consistently yields sharper estimations than the <inline-formula><mml:math id="M94" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, whilst the perceptual loss function is more artifact-prone and thus less predictable. <inline-formula><mml:math id="M95" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> is a common loss function in many image regression tasks, as in image-to-image translation [<xref ref-type="bibr" rid="CIT176">Iso+17</xref>] or in image synthesis [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>].</p>
<fig id="fig-2.20">
<label>Figure 2.20:</label>
<caption><title>Impact of the choice of the loss function used to train the networks in the estimated maps at macroscale. On the left, input guidance X and visual attribute <italic>w</italic> to be transferred. On their right, the result of <inline-formula><mml:math id="M96" display='block'><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>&#x03C9;</mml:mi></mml:msub></mml:math></inline-formula> when trained using <inline-formula><mml:math id="M97" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>, <inline-formula><mml:math id="M98" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> and <italic>LPIPS</italic> [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] loss functions. All the networks in this experiment were trained with the same configuration, only the loss function was modified.</title></caption>
<alt-text>Impact of the choice of the loss function used to train the networks in the estimated maps at macroscale. On the left, input guidance X and visual attribute w to be transferred. On their right, the result of M&#x03C9; when trained using L1, L2 and LPIPS [Zha+18] loss functions. All the networks in this experiment were trained with the same configuration, only the loss function was modified.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.20.jpg"/>
</fig>
</sec>
<sec id="c2-sA.4">
<title>Additional Dataset Details</title>
<p>We show the positions of the lights in the photometric dataset in <xref ref-type="fig" rid="fig-2.21">Figure 2.21</xref>.</p>
<fig id="fig-2.21">
<label>Figure 2.21:</label>
<caption><title>Relative position of the 27 directional light sources used to capture the photometric datasets with respect to the material sample, in degrees. In blue, we show the three lights that are always included in our reduced photometric datasets for the light sampling experiments. As shown, the lights cover a wide variety of azimuth and zenith angles, allowing for an accurate capture of the optical behavior of the materials.</title></caption>
<alt-text>Relative position of the 27 directional light sources used to capture the photometric datasets with respect to the material sample, in degrees. In blue, we show the three lights that are always included in our reduced photometric datasets for the light sampling experiments. As shown, the lights cover a wide variety of azimuth and zenith angles, allowing for an accurate capture of the optical behavior of the materials.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.21.jpg"/>
</fig>
</sec>
</sec>
<sec id="c2-sB">
<label>2.B</label>
<title>Additional Comparisons</title>
<sec id="c2-sB.1">
<title>Comparison with Image Analogies Methods</title>
<p>In <xref ref-type="fig" rid="fig-2.22">Figure 2.22</xref>, we compare our method to different image analogies approaches. We show a comparison between [<xref ref-type="bibr" rid="CIT150">Her+01</xref>], [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>], [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>], and our method. The seminal work in image analogies [<xref ref-type="bibr" rid="CIT150">Her+01</xref>] fails in this visual attribute transfer problem, most likely due to the difference in scale between input and target images. For the comparisons with [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>] and [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>], we reduced <target target-type="page" id="pges_55"/>the scale of the target images <italic>D</italic> to 256 &#x00D7; 256 pixels to make their methods computationally tractable. The methods in [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>; <xref ref-type="bibr" rid="CIT025">Ben+21</xref>] never see the diffuse image A, as they generate an intermediate representation by themselves. While showing better results than [<xref ref-type="bibr" rid="CIT150">Her+01</xref>], methods which use deep learning for this task still fail to yield compelling or predictable results. The method in [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>], which leverages structural information in the images to train a single-image generative model [<xref ref-type="bibr" rid="CIT359">SDM19</xref>], shows more plausible results, but with an expensive computational footprint (10 hours per network), compared to the 45 minutes in [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>], 10 minutes in [<xref ref-type="bibr" rid="CIT150">Her+01</xref>] and less than 1 minute in our case. Because we assume pixel-wise semantic correspondence between input image and visual attributes, we can find mappings in interactive times, and which are significantly more predictable and of higher quality than these image analogies methods.</p>
<fig id="fig-2.22">
<label>Figure 2.22:</label>
<caption><title>Comparison of photometricNet with related image analogies methods. In order, we show a comparison between [<xref ref-type="bibr" rid="CIT150">Her+01</xref>], [<xref ref-type="bibr" rid="CIT245">Lia+17</xref>], [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>], and our method. Our method clearly outperforms any analogies method, at a fraction of the computational cost. The comparisons with [<xref ref-type="bibr" rid="CIT150">Her+01</xref>] were done using an unofficial implementation of their method.</title></caption>
<alt-text>Comparison of photometricNet with related image analogies methods. In order, we show a comparison between [Her+01], [Lia+17], [Ben+21], and our method. Our method clearly outperforms any analogies method, at a fraction of the computational cost. The comparisons with [Her+01] were done using an unofficial implementation of their method.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.22.jpg"/>
</fig>
</sec>
<sec id="c2-sB.2">
<title><target target-type="page" id="pges_56"/>Comparison with Texler et al.</title>
<p>In <xref ref-type="fig" rid="fig-2.23">Figure 2.23</xref>, we show a comparison with different configurations of [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>]. Specifically, we train their method using a dataset with a single image using diffuse illumination (left) and the full photometric dataset (middle column). Their method is trained individually for each dataset and visual attribute, for around 5 minutes to guarantee convergence, as suggested in their paper. As [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>] assumes that the scale (pixels/cm) of the dataset and evaluation images are the same, we rescale our input dataset to match the scale of the evaluation images. We evaluate each method using three different images X, one with diffuse-light illumination (top row), another with diffuse-light illumination but with a geometric transformation (middle row), and an image captured with directional lighting (bottom row). As shown, by enhancing the method in [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>] with a photometric dataset, their model is able to generalize better to every image X in the evaluation dataset than its counterpart trained only using a single image. However, our method, specifically <target target-type="page" id="pges_57"/>designed for this illumination and geometry invariant visual attribute transfer outperforms their work in every type of input and visual attribute, whilst maintaining a lower computational overhead (1 minute of training in our case) and memory footprint (their trained network uses up 12 Mbs, while ours is 1.9 Mbs). Our model trained using a single illumination (diffuseNet) shows better results than their method trained using the photometric dataset, suggesting that our transfer method has more adequate inductive biases for the visual attribute transfer problem.</p>
<fig id="fig-2.23">
<label>Figure 2.23:</label>
<caption><title>Comparison of our method with different configurations of [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>]. On the top row, we compare their results and ours using a dataset containing only one image, captured with diffuse lighting. On the bottom row, we show results when training these methods with a full photometric dataset. Our method is able to generalize better to different illumination and geometric conditions, with a smaller computational overhead.</title></caption>
<alt-text>Comparison of our method with different configurations of [Tex+20b]. On the top row, we compare their results and ours using a dataset containing only one image, captured with diffuse lighting. On the bottom row, we show results when training these methods with a full photometric dataset. Our method is able to generalize better to different illumination and geometric conditions, with a smaller computational overhead.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.23.jpg"/>
</fig>
</sec>
<sec id="c2-sB.3">
<title>Comparison with deep SVBRDF acquisition</title>
<p>In <xref ref-type="fig" rid="fig-2.24">Figure 2.24</xref>, we show a comparison with a recent deep SVBRDF acquisition method [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>], which relies on a pre-trained autoencoder and a differentiable render engine for optimizing the initial estimation of the autoencoder. As shown, our method generates sharper and more coherent normal maps. PhotometricNet relies on a small sample of the same material for training, and is thus less sensitive to potential biases of the training dataset, with the cost of having reduced generalization capabilities compared to multi-material methods.<target target-type="page" id="pges_58"/><target target-type="page" id="pges_59"/></p>
<fig id="fig-2.24">
<label><target target-type="page" id="pges_60"/>Figure 2.24:</label>
<caption><title>Comparison of our method with [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>]. From left to right, we show the input guidance X, the output of photometricNet <inline-formula><mml:math id="M99" display='block'><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>&#x03C9;</mml:mi></mml:msub></mml:math></inline-formula> trained to transfer normal maps, and the output of [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>] after 5000 optimization steps. Their method relies on a pre-trained network, which learns a prior over materials that does not generalize to examples outside their training set. The images in this experiment are of 256 &#x00D7; 256 pixels, to match the input requirements in [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>].</title></caption>
<alt-text>Comparison of our method with [Gao+19]. From left to right, we show the input guidance X, the output of photometricNet M&#x03C9; trained to transfer normal maps, and the output of [Gao+19] after 5000 optimization steps. Their method relies on a pre-trained network, which learns a prior over materials that does not generalize to examples outside their training set. The images in this experiment are of 256 &#x00D7; 256 pixels, to match the input requirements in [Gao+19].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.24.jpg"/>
</fig>
</sec>
</sec>
<sec id="c2-sC">
<label>2.C</label>
<title>Extension to SVBSDF Propagation</title>
<p>In this section, we present an extension to the previous method for propagating material maps. With it, we can propagate multiple property maps at the same time, learn from many photometric datasets simultaneously, and achieve more accurate <target target-type="page" id="pges_61"/>and sharper results. With these improvements, we are able to propagate multiple SVBSDF maps at the same time, allowing for efficient high-resolution material generation, as we show in <xref ref-type="fig" rid="fig-2.25">Figures 2.25</xref> and <xref ref-type="fig" rid="fig-2.26">2.26</xref>. These findings are part of the data generation process which is used to build the datasets we use to train the models described in <xref ref-type="sec" rid="c2-s5">Section 5</xref> and [<xref ref-type="bibr" rid="CIT337">Rod+23a</xref>], and are also part of a paper on high resolution material digitization, which has been accepted to <italic>SIGGRAPH 2023</italic>:</p>
<fig id="fig-2.25">
<label>Figure 2.25:</label>
<caption><title>Renders of SVBSDF maps transferred by our extension of photometricNet.</title></caption>
<alt-text>Renders of SVBSDF maps transferred by our extension of photometricNet.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.25.jpg"/>
</fig>
<fig id="fig-2.26">
<label>Figure 2.26:</label>
<caption><title>Left, virtual scene rendered using eight different materials digitized with our optical device. Right, real photos of four of the materials taken under diverse illumination conditions: area lights, diffuse lighting, and directional lighting of a high-resolution patch.</title></caption>
<alt-text>Left, virtual scene rendered using eight different materials digitized with our optical device. Right, real photos of four of the materials taken under diverse illumination conditions: area lights, diffuse lighting, and directional lighting of a high-resolution patch.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.26.jpg"/>
</fig>
<fig id="fig-2">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.jpg"/>
</fig>
<disp-quote>
<p>&#x201C;Towards Material Digitization with a Dual-scale Optical System&#x201D; Elena Garces, Victor Arellano, Carlos Rodriguez-Pardo, David Pascual-Hernandez, Sergio Suja, Jorge Lopez-Moreno</p>
<p><italic>ACM Transactions on Graphics (Proceedings of SIGGRAPH)</italic> (2023)</p>
</disp-quote>
<sec id="c2-sC.s1">
<label>2.C.1</label>
<title>Mesoscale Propagation</title>
<p>To propagate material properties, a common solution is to utilize the estimation of material properties in a localized area and extended it to larger samples [<xref ref-type="bibr" rid="CIT294">ND06</xref>; <xref ref-type="bibr" rid="CIT345">RSK13</xref>; <xref ref-type="bibr" rid="CIT374">Ste+14</xref>; <xref ref-type="bibr" rid="CIT067">DDB20</xref>]. We follow this approach to propagate SVBSDF measurements, and build upon the work presented in <xref ref-type="sec" rid="c2-s1">Section 2</xref> and [<xref ref-type="bibr" rid="CIT338">RG21</xref>], <target target-type="page" id="pges_62"/>which leverage images of the material taken under different illumination conditions to train a neural network that propagates individual maps.</p>
<p>As input to this propagation step, we use a photometric dataset obtained with a microscopic setup, spatially-varying parameters of a micro SVBSDF (Illustrated in <xref ref-type="fig" rid="fig-2.28">Figure 2.28</xref>), and a guidance image that will serve as a reference for the mesoscale propagation. We use as guidance the image taken with a mid-range camera and diffuse lighting. Unlike the previous approach, aimed at transferring individual maps, we introduce some modifications to jointly propagate the full stack of parameters. We further adapt the method to be able to take as input more than one micro SVBSDF, necessary if the variability of material is not visible in a single micro capture (see <xref ref-type="fig" rid="fig-2.27">Figure 2.27</xref> (a)). We propose the following modifications:</p>
<list list-type="order">
<list-item><p><bold>Larger Training Dataset.</bold> We improve the generalization of the model by using more training data in two ways: First, we use a larger photometric dataset, using more directional-lit images than what is described in 2.21, as well as diffuse lighting. Second, we use random Gaussian blurs for data augmentation to improve robustness to potential degradation in mesoscale guidance images, and remove affine distortions to the data augmentation policy as they are not required for our SVBSDF transfer problem. Additionally, we train the model for more iterations, with a larger batch size.</p></list-item>
<list-item><p><bold>Enable Large Spatially-varying Materials.</bold> We account for materials in which the appearance variability is not observed in a single micro capture by training the propagation network with as many micro SVBSDF as necessary. <xref ref-type="fig" rid="fig-2.27">Figure 2.27</xref> (a) showcases an example that required two microscale captures. Further, we disable the color-invariance data augmentation that randomly shuffles color channels when multiple micro SVBSDFs are available.</p></list-item>
<list-item><p><bold>Improved Architecture and Losses.</bold> Finally, we observed that the original architecture was not able to effectively transfer a large number of parameters at the same time. Therefore, inspired by recent work on intrinsic decomposition [<xref ref-type="bibr" rid="CIT179">Jan+17</xref>], texture synthesis [<xref ref-type="bibr" rid="CIT339">RG22</xref>], and material estimation [<xref ref-type="bibr" rid="CIT068">DLG21</xref>], we use a U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] with a single encoder but a separate decoder for every map we aim to transfer. Further, we introduce residual connections [<xref ref-type="bibr" rid="CIT139">He+16</xref>; <xref ref-type="bibr" rid="CIT069">Dia+20</xref>] into every layer, and substitute Batch Normalization [<xref ref-type="bibr" rid="CIT175">IS15</xref>] for Group Normalization [<xref ref-type="bibr" rid="CIT427">WH18</xref>]. When multiple microscale captures are available, we use a model with additional filters in every layer, to better learn from these extended datasets. This design generates sharper texture maps which preserve the statistics and appearance of each fitted microscale map more accurately. We also added a multi-channel perceptual component [<xref ref-type="bibr" rid="CIT041">CHB21</xref>] to the loss function, which is a powerful <target target-type="page" id="pges_63"/>regularizer for texture synthesis. <xref ref-type="fig" rid="fig-2.27">Figure 2.27</xref> (b-c) shows the difference with respect to the previous approach.</p></list-item>
</list>
<fig id="fig-2.27">
<label><target target-type="page" id="pges_64"/>Figure 2.27:</label>
<caption><title>Comparison with photometricNet [<xref ref-type="bibr" rid="CIT338">RG21</xref>] for mesoscale maps propagation. The first row (a) showcases an example where multiple captures at microscale were needed to cover the spatially-varying albedo of the material. (b) and (c) required a single capture, we show the result of normals and tangents maps compared with previous work.</title></caption>
<alt-text>Comparison with photometricNet [RG21] for mesoscale maps propagation. The first row (a) showcases an example where multiple captures at microscale were needed to cover the spatially-varying albedo of the material. (b) and (c) required a single capture, we show the result of normals and tangents maps compared with previous work.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.27.jpg"/>
</fig>
<fig id="fig-2.28">
<label>Figure 2.28:</label>
<caption><title>On the left, microscale SVBSDF maps. On the right, a render generated with the maps propagated with our method.</title></caption>
<alt-text>On the left, microscale SVBSDF maps. On the right, a render generated with the maps propagated with our method.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.28.jpg"/>
</fig>
</sec>
<sec id="c2-sC.s2">
<label>2.C.2</label>
<title>Implementation Details</title>
<sec id="c2-sC.s2.1">
<title>Problem Formulation</title>
<p>As in [<xref ref-type="bibr" rid="CIT338">RG21</xref>], our goal is to transfer spatially-varying attributes of a material captured at microscale to larger samples of the same material. We thus formulate this problem using an image-to-image translation framework. For training, our methods take as input a <italic>set of D photometric datasets</italic> of a material <inline-formula><mml:math id="M100" display='block'><mml:msubsup><mml:mi mathvariant="script">I</mml:mi><mml:mi>L</mml:mi><mml:mi>D</mml:mi></mml:msubsup></mml:math></inline-formula>, and their pixel-wise corresponding SVBSDFs <inline-formula><mml:math id="M101" display='block'><mml:msubsup><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi></mml:msubsup></mml:math></inline-formula>, each comprised of a list of <italic>p</italic> maps: <italic>p</italic> &#x2208; {albedo, roughness, transmittance, IOR, anisotropy, tangents, normals, specularTint, opacity}. Each photometric dataset is comprised of a set of RGB images of the material captured under different illumination conditions: <inline-formula><mml:math id="M102" display='block'><mml:msubsup><mml:mi mathvariant="script">I</mml:mi><mml:mi>L</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>&#x2223;</mml:mo><mml:msubsup><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo></mml:mrow><mml:mo>&#x00A0;</mml:mo><mml:mrow><mml:msup><mml:mi mathvariant="fraktur">R</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mi mathvariant="script">I</mml:mi><mml:mi>L</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>D</mml:mi></mml:math></inline-formula>. In contrast to [<xref ref-type="bibr" rid="CIT338">RG21</xref>], which only allowed to learn from a single photometric dataset and to transfer a single attribute, our extended approach allows for learning from multiple datasets |<italic>D</italic>| &#x2265; 1 and multiple property maps |<italic>p</italic> | &#x2265; 1 using a single model. We train a model <italic>T</italic> which learns to transfer from each image in each photometric dataset to its corresponding SVBSDF: <inline-formula><mml:math id="M103" display='block'><mml:mi>T</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2248;</mml:mo><mml:msubsup><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>D</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula>. During inference, we use a <inline-formula><mml:math id="M104" display='block'><mml:mi mathvariant="bold-italic">l</mml:mi></mml:math></inline-formula> <italic>p</italic> guidance image of the material X, which represents a larger sample of it, and estimate its corresponding property maps: <inline-formula><mml:math id="M105" display='block'><mml:msup><mml:mi>M</mml:mi><mml:mi mathvariant="normal">X</mml:mi></mml:msup><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mi>T</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">X</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula>. At evaluation time, this guidance image X may be captured with different conditions to those in the training dataset, including different camera or illumination conditions.</p>
</sec>
<sec id="c2-sC.s2.2">
<title>Datasets</title>
<p>We use a more comprehensive photometric dataset to what is proposed in [<xref ref-type="bibr" rid="CIT338">RG21</xref>]. In particular, for a single microscale capture, we use all the light sources used in the fitting algorithm, as well as diffuse-lit microscale images. Additionally, we use both polarization modes separately and, for the same light source, we construct an extra image by averaging the images taken under both polarization configurations. The total size of each photometric dataset is of <inline-formula><mml:math id="M106" display='block'><mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="script">I</mml:mi></mml:mrow><mml:mi>L</mml:mi><mml:mi>d</mml:mi></mml:msubsup><mml:mo>|</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>108</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>D</mml:mi></mml:math></inline-formula>, four times more data compared with the more limited 27 images proposed in [<xref ref-type="bibr" rid="CIT338">RG21</xref>]. For evaluation, our guidance image is taken with a mid-range camera and <italic>diffuse</italic> illumination, at a resolution of 4072 &#x00D7; 4072 pixels and a surface area of 10 &#x00D7; 10 centimeters. We typically use a single photometric dataset (<inline-formula><mml:math id="M107" display='block'><mml:mo stretchy="false">|</mml:mo><mml:mi>D</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>). However, for multi-colored or <target target-type="page" id="pges_65"/>highly heterogeneous materials in which a single capture is not enough to fully represent the material variability, we capture as many datasets as needed.</p>
</sec>
<sec id="c2-sC.s2.3">
<title>Loss Function</title>
<p>Our loss function is the combination of two terms: weighted pixel-wise losses for each target map, and a multi-map style loss:</p>
<p><disp-formula id="Eq002-3"><label>(2.3)</label> <mml:math id="M108" display='block'><mml:mi mathvariant="script">L</mml:mi><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>p</mml:mi></mml:munder><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mtext>pixel</mml:mtext><mml:mi>p</mml:mi></mml:msub></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>style</mml:mtext></mml:msub></mml:math></disp-formula></p>
<p><inline-formula><mml:math id="M109" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>pixel</mml:mtext></mml:msub></mml:math></inline-formula> is the <inline-formula><mml:math id="M110" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> norm weighted per map in the SVBSDF, <inline-formula><mml:math id="M111" display='block'><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:math></inline-formula>. <inline-formula><mml:math id="M112" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> produces sharper results than higher-order alternatives [<xref ref-type="bibr" rid="CIT338">RG21</xref>], such as <inline-formula><mml:math id="M113" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>. For the <inline-formula><mml:math id="M114" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>style</mml:mtext></mml:msub></mml:math></inline-formula> loss term, we follow recent work on multi-channel texture synthesis, which extends deep perceptual losses for style transfer to the SVBRDF synthesis problem [<xref ref-type="bibr" rid="CIT041">CHB21</xref>]. This component of the loss function acts as a regularizer and allows the model to generate maps that better preserve the appearance of the ground truth training data.</p>
</sec>
<sec id="c2-sC.s2.4">
<title>Training Implementation Details</title>
<p><bold>Model Design</bold> We use a lightweight U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] architecture, with a few modifications based on recent work to maximize its efficiency and the quality of its outputs. We specify the full model architecture and model sizes in <xref ref-type="fig" rid="fig-2.29">Figure 2.29</xref>. In every convolutional block of the model, we use residual connections [<xref ref-type="bibr" rid="CIT139">He+16</xref>; <xref ref-type="bibr" rid="CIT069">Dia+20</xref>], for better training convergence and preservation of details in the input images. We use 1 &#x00D7; 1 convolutions on these residual connections. We use a single decoder for each map, so as to maximally preserve their individual appearance and statistics. This has been proposed in recent work on texture synthesis [<xref ref-type="bibr" rid="CIT339">RG22</xref>], material capture [<xref ref-type="bibr" rid="CIT068">DLG21</xref>] or intrinsic decomposition [<xref ref-type="bibr" rid="CIT179">Jan+17</xref>]. On the convolutional blocks, we leverage Group Normalization [<xref ref-type="bibr" rid="CIT427">WH18</xref>]. Upsampling is done using transposed convolutions. As shown in <xref ref-type="fig" rid="fig-2.29">Figure 2.29</xref>, the model has four hidden layers on the encoder and each decoder, however, we vary the size of those layers <target target-type="page" id="pges_66"/>depending on whether the model is trained with a single microscale capture (for which we use a width factor of <italic>W</italic> = 16) or multiple (<italic>W</italic> = 32). Every other implementation detail in the model (stride, bias, pooling) follows [<xref ref-type="bibr" rid="CIT344">RFB15</xref>; <xref ref-type="bibr" rid="CIT338">RG21</xref>].</p>
<fig id="fig-2.29">
<label><target target-type="page" id="pges_67"/>Figure 2.29:</label>
<caption><title>A full diagram of the model architecture we use for our mesoscale propagation problem. Building upon [<xref ref-type="bibr" rid="CIT338">RG21</xref>], use a lightweight U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] architecture. Following previous work [<xref ref-type="bibr" rid="CIT179">Jan+17</xref>; <xref ref-type="bibr" rid="CIT339">RG22</xref>; <xref ref-type="bibr" rid="CIT068">DLG21</xref>], we use a different decoder for every map we aim to transfer. Each of our convolutional blocks, illustrated on the right, contain residual connections [<xref ref-type="bibr" rid="CIT139">He+16</xref>; <xref ref-type="bibr" rid="CIT069">Dia+20</xref>] and Group Normalization [<xref ref-type="bibr" rid="CIT427">WH18</xref>]. In red, we show the input/output dimensions (spatial, channels) of each layer; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, regularizations, and non-linearities. We set <italic>W</italic> to <italic>W</italic> = 16 when a single microscale capture is available as training data, <italic>W</italic> = 32 otherwise.</title></caption>
<alt-text>A full diagram of the model architecture we use for our mesoscale propagation problem. Building upon [RG21], use a lightweight U-Net [RFB15] architecture. Following previous work [Jan+17; RG22; DLG21], we use a different decoder for every map we aim to transfer. Each of our convolutional blocks, illustrated on the right, contain residual connections [He+16; Dia+20] and Group Normalization [WH18]. In red, we show the input/output dimensions (spatial, channels) of each layer; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, regularizations, and non-linearities. We set W to W = 16 when a single microscale capture is available as training data, W = 32 otherwise.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.29.jpg"/>
</fig>
<p><bold>Training and Evaluation</bold> We use PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>], Torchvision [<xref ref-type="bibr" rid="CIT262">MR10</xref>], and Kornia [<xref ref-type="bibr" rid="CIT334">Rib+20</xref>] for training. We leverage mixed precision training and automatic gradient scaling [<xref ref-type="bibr" rid="CIT274">Mic+18</xref>], to accelerate the training process and regularize the models. Optimization is done using Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>] for 5000 iterations, with a learning rate of 0.002, a batch size of 40, and no weight decay. This process takes around 15 minutes when training with a single microscale capture, and 25 when using multiple captures as input. We evaluate the guidances images using half precision. We measure these times on an NVIDIA 1080Ti GPU.</p>
<p><bold>Data Augmentation</bold> We follow the same training procedure defined in [<xref ref-type="bibr" rid="CIT338">RG21</xref>]. However, we do not use random rotations or shears, and remove the color invariance data augmentation whenever multiple microscale captures are available. Further, we introduce random Gaussian Blurs for data augmentation, using a <italic>p</italic> = 0.5, a kernel size of 5, and sigma selected uniformly at random for each element in each batch: <inline-formula><mml:math id="M115" display='block'><mml:mi>&#x03C3;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>U</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>11</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. As in [<xref ref-type="bibr" rid="CIT338">RG21</xref>], we also perform random cropping, using crops of 128 &#x00D7; 128 pixels. For each element in each batch during training, we randomly choose the photometric dataset d &#x2208; <italic>D</italic>, and, from it, a random light source <italic>l</italic> &#x2208; <italic>L</italic>.</p>
<p><bold>Loss Function</bold> We weight the loss function as follows: For the pixel wise loss: <inline-formula><mml:math id="M117" display='block'><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>normals</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mi>O</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>roughness</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>albedo</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>tangent</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>anisotropy</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>specularTint</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>transmittance</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>opacity</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. The perceptual component [<xref ref-type="bibr" rid="CIT041">CHB21</xref>] is weighted with <inline-formula><mml:math id="M118" display='block'><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>style </mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.25</mml:mn></mml:math></inline-formula>, and we use the <italic>AlexNet</italic> variant of <italic>LPIPS</italic> [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] as our backbone, which provides a powerful yet lightweight loss for texture transfer, as shown in [<xref ref-type="bibr" rid="CIT338">RG21</xref>].</p>
</sec>
</sec>
<sec id="c2-sC.s3">
<label>2.C.3</label>
<title>Results</title>
<p>In <xref ref-type="fig" rid="fig-2.30">Figure 2.30</xref>, we show an ablation study on the improvements of the maps propagation method proposed in [<xref ref-type="bibr" rid="CIT338">RG21</xref>]. We build upon their proposed implementation, removing their random shifts and rotations, and make progressive changes, in order, to the dataset size, model architecture, loss function, and data augmentation. We observe high-quality propagations, but the baseline sometimes produces overly smooth outputs (<sc>Fleece, Leather-Brown, Rib-Silver, or Leather-Lizard</sc>). Our proposed increased dataset and model architecture modifications enhance the maps&#x2019; sharpness, and our perceptual loss function allows for the propagation of additional details. Our final model removes the color augmentation proposed in [<xref ref-type="bibr" rid="CIT338">RG21</xref>] whenever multiple microscale captures are available, and introduces random blurs during training, allowing for achieving the highest quality results and for accurate base color propagations in challenging cases, as in <sc>Tartan</sc>.</p>
<fig id="fig-2.30">
<label><target target-type="page" id="pges_68"/>Figure 2.30:</label>
<caption><title>Ablation study of our proposed improvements with respect to tphotometricNet [<xref ref-type="bibr" rid="CIT338">RG21</xref>], on base color (top rows), normals (mid) and tangent map (bottom) transfer. We start from the implementation in [<xref ref-type="bibr" rid="CIT338">RG21</xref>], which shows adequate but smooth results. On its right, we show the impact of increasing the dataset size, which tends to generate higher-quality estimations (Leather-Brown). With our improved architecture (fifth column), we achieve sharper maps (see Leather-Lizard, Rib-Silver, or Fleece). Using a perceptual loss (sixth column), we further improve the maps quality (Leather-Brown). Our final model (last column), which removes the color augmentation in [<xref ref-type="bibr" rid="CIT338">RG21</xref>] and introduces random blurs to the model training, achieves the highest quality, and allows for accurate base color transfers, as shown in Tartan. Input guidance images and training micro maps are shown on the first and second columns, respectively.</title></caption>
<alt-text>Ablation study of our proposed improvements with respect to tphotometricNet [RG21], on base color (top rows), normals (mid) and tangent map (bottom) transfer. We start from the implementation in [RG21], which shows adequate but smooth results. On its right, we show the impact of increasing the dataset size, which tends to generate higher-quality estimations (Leather-Brown). With our improved architecture (fifth column), we achieve sharper maps (see Leather-Lizard, Rib-Silver, or Fleece). Using a perceptual loss (sixth column), we further improve the maps quality (Leather-Brown). Our final model (last column), which removes the color augmentation in [RG21] and introduces random blurs to the model training, achieves the highest quality, and allows for accurate base color transfers, as shown in Tartan. Input guidance images and training micro maps are shown on the first and second columns, respectively.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-2.30.jpg"/>
</fig>
</sec>
</sec>
</body>
<back>
<fn-group>
<fn id="c2-fn1"><label>1</label> <p>Additional publication details and results are included on <bold>the project website</bold>.</p></fn>
</fn-group>
</back>
</book-part>
<book-part id="c3" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_69"/>Chapter 3.</label>
<title>A Single Image Generative Model for Tileable Texture Synthesis</title>
</title-group>
</book-part-meta>
<body>
<fig id="fig-3.1">
<label>Figure 3.1:</label>
<caption><title>Results obtained with SeamlessGAN for tileable texture synthesis. We show 2x2 repetitions of the outputs on the squared images, and the input exemplars by their side.</title></caption>
<alt-text>Results obtained with SeamlessGAN for tileable texture synthesis. We show 2x2 repetitions of the outputs on the squared images, and the input exemplars by their side.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.1.jpg"/>
</fig>
<p>In this chapter, we present <italic>SeamlessGAN</italic>, a method capable of automatically generating tileable texture maps from a single input exemplar. In contrast to most existing methods, focused solely on solving the synthesis problem, our work tackles both problems, <italic>synthesis</italic> and <italic>tileability</italic>, simultaneously. Our key idea is to realize that tiling a latent space within a generative network trained using adversarial expansion techniques produces outputs with continuity at the seam intersection that can then be turned into tileable images by cropping the central area. Since not every value of the latent space is valid to produce high-quality outputs, we leverage the discriminator as a perceptual error metric capable of identifying artifact-free textures during a sampling process. Further, in contrast to previous work on deep texture synthesis, our model is designed and optimized to work with multi-layered texture representations, enabling textures composed of multiple maps such as albedo, normals, etc. We extensively test our design choices for the network architecture, loss function, and sampling parameters. We show qualitatively and quantitatively that our approach outperforms previous methods and works for textures of different types. The contributions presented in this chapter have led to the following publication<xref ref-type="fn" rid="c3-fn1"><sup>1</sup></xref>:</p>
<fig id="fig-3">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.jpg"/>
</fig>
<p><target target-type="page" id="pges_70"/>&#x201C;SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps&#x201D; Carlos Rodriguez-Pardo, Elena Garces</p>
<p><italic>IEEE Transactions on Visualization and Computer Graphics (TVCG)</italic></p>
<p>(2022)</p>
<sec id="c3-s1">
<label>3.1</label>
<title>Introduction</title>
<p>Realistic and high-quality textures are important elements to convey realism in virtual environments. These can be procedurally generated [<xref ref-type="bibr" rid="CIT392">Tu+20</xref>; <xref ref-type="bibr" rid="CIT163">HDR19</xref>; <xref ref-type="bibr" rid="CIT123">Gue+20</xref>; <xref ref-type="bibr" rid="CIT098">Gal+12</xref>; <xref ref-type="bibr" rid="CIT115">Gil+14</xref>; <xref ref-type="bibr" rid="CIT124">Gui+17</xref>; <xref ref-type="bibr" rid="CIT144">HN18</xref>], captured [<xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT242">LSC18</xref>] or synthesized from real images [<xref ref-type="bibr" rid="CIT083">EF01</xref>; <xref ref-type="bibr" rid="CIT222">Kwa+05</xref>; <xref ref-type="bibr" rid="CIT458">Zho+18</xref>; <xref ref-type="bibr" rid="CIT281">Mor+17</xref>]. Frequently, textures are used to efficiently reproduce elements with repetitive patterns (for example, facades, surfaces, or materials) by means of spatially concatenating -or <italic>tiling</italic>- multiple copies of themselves. Creating tileable textures is a very challenging problem, as it requires a semantic understanding of the repetitive elements, often at multiple scales. For this reason, such a process is frequently done manually by artists in 3D digitization pipelines.</p>
<p>Recent advances in Convolutional Neural Networks (CNNs) and Generative Adversarial Networks (GANs) have been applied to texture synthesis problems [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>; <xref ref-type="bibr" rid="CIT095">FAW19</xref>; <xref ref-type="bibr" rid="CIT248">Liu+20</xref>; <xref ref-type="bibr" rid="CIT028">BJV17</xref>; <xref ref-type="bibr" rid="CIT182">JBV16</xref>; <xref ref-type="bibr" rid="CIT263">Mar+20</xref>; <xref ref-type="bibr" rid="CIT149">Her+20</xref>] showing unprecedented levels of realism and quality, however, the output of these methods is not tileable. Despite recent methods [<xref ref-type="bibr" rid="CIT063">DH19</xref>; <xref ref-type="bibr" rid="CIT028">BJV17</xref>; <xref ref-type="bibr" rid="CIT341">Rod+19</xref>; <xref ref-type="bibr" rid="CIT281">Mor+17</xref>; <xref ref-type="bibr" rid="CIT241">Li+20</xref>; <xref ref-type="bibr" rid="CIT296">Nik+21</xref>] addressing the problem of tileable texture synthesis, we show that they either assume a particular level of regularity or the generated textures lose a significant amount of visual fidelity with respect to the input exemplars. Further, most of these methods have only focused on synthesizing single images. Rendering realistic materials requires more information about their optical properties beyond what represents a single RGB pixel. To this end, it is common to use spatially-varying BRDFs [<xref ref-type="bibr" rid="CIT066">Des+19</xref>], which are optical appearance models parameterized by stacks of images, each one representing a different property, such as albedo, normals, or transparency. As the number of methods that generate texture stacks from physical samples grows [<xref ref-type="bibr" rid="CIT066">Des+19</xref>; <xref ref-type="bibr" rid="CIT067">DDB20</xref>; <xref ref-type="bibr" rid="CIT127">Guo+20b</xref>] so does the need to turn them into <italic>tileable texture stacks</italic>.</p>
<p>In this chapter, we propose a deep generative model, <italic>SeamlessGAN</italic>, capable of synthesizing tileable texture stacks from inputs of arbitrary content. In contrast <target target-type="page" id="pges_71"/>to <italic>Wang Tiles</italic> [<xref ref-type="bibr" rid="CIT055">Coh+03</xref>], by which a single large texture region is created by concatenating multiple different tiles with matching borders, we aim to automatically obtain a seamless single-tile, which allows for reduced memory consumption and enhanced usability in casual scenarios when the user might lack the necessary artistic skills. Our key idea is to realize that tiling a latent space within a generative network produces outputs with continuity at the seam intersection [<xref ref-type="bibr" rid="CIT095">FAW19</xref>], which can then be turned into tileable images by cropping the central area. Since not every value of the latent space is valid to produce high-quality outputs, we follow a double strategy: First, we train the generative network using an adversarial expansion technique [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>], which provides a latent encoding of the input texture, which can then be decoded into high-quality outputs that double the spatial extent of the input. Second, we use the trained discriminator as a local quality metric in a sampling algorithm. This allows us to find the input of the generative process that produces tileable textures similar to the original, as well as multiple candidates per input exemplar. As opposed to previous work, which focused on maximizing stationarity [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>] and thus might remove important high-level texture features in regular or near-regular textures, our method is focused on maximizing tileability while preserving the original texture as intact as possible, in terms of its stylistic and semantic properties. To allow for the synthesis of stacks of textures, we propose a neural architecture composed of various decoder networks. Despite not explicitly imposing inter-map consistency, we show that it is implicitly guaranteed by how the generative network is trained. Without losing generality, we show texture stacks synthesis of two maps: an albedo and a normal map. We demonstrate that our method outperforms state-of-the-art solutions on tileable texture synthesis of single images and show several examples for synthesizing tileable texture stacks. We validate our design choices through several ablations studies and off-the-shelf perceptual quality metrics.</p>
<fig id="fig-3.2">
<label>Figure 3.2:</label>
<caption><title>Na&#x00EF;vely tiling the original texture causes discontinuities at the seam intersections as shown in the top row. Our method automatically generates a tileable texture stack from an input exemplar which double the size of the input.</title></caption>
<alt-text>Na&#x00EF;vely tiling the original texture causes discontinuities at the seam intersections as shown in the top row. Our method automatically generates a tileable texture stack from an input exemplar which double the size of the input.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.2.jpg"/>
</fig>
</sec>
<sec id="c3-s2">
<label>3.2</label>
<title>Related work</title>
<p>We review the texture synthesis methods most closely related to our work. For a more comprehensive review, please refer to the surveys in [<xref ref-type="bibr" rid="CIT006">Akl+18</xref>; <xref ref-type="bibr" rid="CIT312">Raa+18</xref>]. We will also mention other related work regarding tileable texture synthesis, and deep internal learning.</p>
<sec id="c3-s2.s1">
<label>3.2.1</label>
<title>Texture Synthesis</title>
<p>Traditionally, <bold>non-parametric texture synthesis</bold> algorithms worked by ensuring that every patch in the output textures approximates a patch in the input texture. Earlier methods included image quilting [<xref ref-type="bibr" rid="CIT083">EF01</xref>; <xref ref-type="bibr" rid="CIT084">EL99</xref>], GraphCuts [<xref ref-type="bibr" rid="CIT223">Kwa+03</xref>], genetic algorithms [<xref ref-type="bibr" rid="CIT075">DZP07</xref>], and optimization [<xref ref-type="bibr" rid="CIT222">Kwa+05</xref>; <xref ref-type="bibr" rid="CIT309">PS00</xref>]. More recent approaches use variations of PatchMatch [<xref ref-type="bibr" rid="CIT022">Bar+09</xref>; <xref ref-type="bibr" rid="CIT023">Bar+15</xref>] as a way of <target target-type="page" id="pges_72"/>finding correspondences between generated and input images [<xref ref-type="bibr" rid="CIT198">Kas+15</xref>; <xref ref-type="bibr" rid="CIT059">Dar+12</xref>; <xref ref-type="bibr" rid="CIT457">Zho+17</xref>].</p>
<p>Despite those methods showing high-quality results for textures of different characteristics, recent work on deep parametric texture synthesis shows better generality and scalability, requiring less manual input. Our approach belongs to the category of <bold>Parametric texture synthesis</bold>. These methods learn statistics from the example textures, which can then be used for the generation of new images that match those statistics. While traditional methods used hand-crafted features [<xref ref-type="bibr" rid="CIT061">De 97</xref>; <xref ref-type="bibr" rid="CIT142">HB95</xref>], recent parametric methods rely on deep neural networks as their parameterization. Activations within latent spaces in pre-trained CNNs have been shown to capture relevant statistics of the style and texture of images [<xref ref-type="bibr" rid="CIT109">GEB16a</xref>; <xref ref-type="bibr" rid="CIT110">Gat+17</xref>; <xref ref-type="bibr" rid="CIT186">JAF16</xref>; <xref ref-type="bibr" rid="CIT330">Ren+17</xref>; <xref ref-type="bibr" rid="CIT342">RB19</xref>]. Textures can be synthesized through this approach by gradient-descent optimization [<xref ref-type="bibr" rid="CIT369">Sne17</xref>; <xref ref-type="bibr" rid="CIT107">GEB15a</xref>] or by training a neural network that learns those features [<xref ref-type="bibr" rid="CIT078">DB16</xref>; <xref ref-type="bibr" rid="CIT296">Nik+21</xref>]. Finding generic patterns that precisely describe the example textures is one of the main challenges in parametric texture synthesis. Features that describe textured images in a generic way are hard to find and they typically require hand-tuning. Generative Adversarial Networks (GANs), which have shown remarkable capabilities in image generation in multiple domains [<xref ref-type="bibr" rid="CIT193">Kar+18</xref>; <xref ref-type="bibr" rid="CIT196">KLA19</xref>; <xref ref-type="bibr" rid="CIT197">Kar+20b</xref>], can learn those features from data. Specifically, in texture synthesis, they have proven successful at generating new samples <target target-type="page" id="pges_73"/>of textures from a single input image [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>] or from a dataset of images [<xref ref-type="bibr" rid="CIT028">BJV17</xref>; <xref ref-type="bibr" rid="CIT095">FAW19</xref>; <xref ref-type="bibr" rid="CIT182">JBV16</xref>; <xref ref-type="bibr" rid="CIT248">Liu+20</xref>; <xref ref-type="bibr" rid="CIT263">Mar+20</xref>]. We build upon the method of Zhou <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>] which shows good performance on the synthesis of non-stationary single image textures, and extend it to synthesize texture stacks, as well as generate tileable outputs.</p>
</sec>
<sec id="c3-s2.s2">
<label>3.2.2</label>
<title>Texture Tileability</title>
<p>Synthesizing tileable textures received surprisingly little attention in the literature until recent years. Moritz <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>] propose a non-parametric approach that is able to synthesize textures from a single example while preserving its <italic>stationarity</italic>, which measures how tileable the texture is. Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT241">Li+20</xref>] propose a GraphCuts-based algorithm. They first find a patch that optimally represents the texture, then use graph cuts to transform its borders to improve its tileability. This method allows for synthesizing multiple maps at the same time. Relatedly, Deliot and Heitz [<xref ref-type="bibr" rid="CIT063">DH19</xref>] propose a histogram-preserving blending operation for patch-based synthesis of tileable textures, particularly suited for stochastic textures. The power of deep neural networks for tileable texture synthesis has also been leveraged in the past years. First, Rodriguez-Pardo <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>] exploit latent spaces in a pre-trained neural network to find the size of the repeating pattern in the input texture. Then, they use perceptual losses for finding the optimal crop of the image such that, when tiled, looks the most similar to the original image. Also leveraging deep perceptual losses and using Neural Cellular Automata [<xref ref-type="bibr" rid="CIT279">Mor+20</xref>] as an image parameterization, Niklasson <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT296">Nik+21</xref>] generate self-organizing textures that are seamlessly tileable by design, but are limited in resolution and by the quality of the gram matrix perceptual metric used as a loss function [<xref ref-type="bibr" rid="CIT145">Hei+21</xref>]. Bergmann <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT028">BJV17</xref>] achieve tileability in their output textures by training a multi-image GAN in a periodic spatial manifold. Our proposed method does not follow any of these approaches. Instead, we build upon a state-of-the-art single-image texture synthesis method, which is able to generate high-quality and high-resolution images, and extend it to generate textures which are tileable. To this end, we propose a sampling algorithm that finds the input of the generative process that maximizes a novel tileability metric. Our goal is to preserve the input original texture as intact as possible while imposing artifact-free continuity at the intersection of the seams when the texture is tiled. Furthermore, our method has a reduced computational footprint compared to other deep generative texture synthesis methods, as we show in the results section.</p>
</sec>
<sec id="c3-s2.s3">
<label>3.2.3</label>
<title>Deep Internal Learning</title>
<p>Learning patterns from a single image has been studied in recent years, in contexts different to those of texture synthesis. Ulyanov <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT395">UVL18</xref>] show <target target-type="page" id="pges_74"/>that single images can be represented by randomly initialized CNNs, and show applicability on denoising or image inpainting problems. A similar method is proposed by Shocher <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT362">SCI18</xref>] for single-image super-resolution. Additionally, single images have shown to be enough for learning low-level features that generalize to multiple problems [<xref ref-type="bibr" rid="CIT014">ARV19</xref>]. Hypernetworks [<xref ref-type="bibr" rid="CIT129">HDL16</xref>] in <italic>Implicit Neural Representations</italic> [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>; <xref ref-type="bibr" rid="CIT384">Tan+20</xref>] also allow for shift-invariant priors over images, which can then used in generative models [<xref ref-type="bibr" rid="CIT079">DTD21</xref>]. Single-image generative models have been explored for domains other than textures. GANs trained on a single image have been used for image retargeting [<xref ref-type="bibr" rid="CIT361">Sho+19</xref>], deep image analogies [<xref ref-type="bibr" rid="CIT025">Ben+21</xref>], or for learning a single-sample image generative model [<xref ref-type="bibr" rid="CIT359">SDM19</xref>; <xref ref-type="bibr" rid="CIT156">Hin+21</xref>; <xref ref-type="bibr" rid="CIT361">Sho+19</xref>; <xref ref-type="bibr" rid="CIT378">SGK21</xref>; <xref ref-type="bibr" rid="CIT402">VZH20</xref>]. These methods, while powerful for natural images, are not well-behaved for textures, as shown in [<xref ref-type="bibr" rid="CIT248">Liu+20</xref>] and [<xref ref-type="bibr" rid="CIT263">Mar+20</xref>]. By introducing inductive biases specially designed for textured images, characterized by their repeating patterns, deep texture synthesis methods achieve better performance in textured images than generic single-image synthesis approaches, which need to account for more globally coherent semantic patterns.</p>
</sec>
</sec>
<sec id="c3-s3">
<label>3.3</label>
<title>Overview</title>
<p>Our method takes as input an <italic>untileable texture stack</italic> <inline-formula><mml:math id="M128" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> , which is a layered representation &#x015B;or SVBRDF [<xref ref-type="bibr" rid="CIT038">BS12</xref>]&#x015B; of a material used by render engines to virtually reproduce it. A typical texture stack is composed of several maps such as albedo, normals, or roughness. For simplicity, and without losing <target target-type="page" id="pges_75"/>generality, we assume that our texture stacks have two maps: an albedo A, and a normal map N, both are RGB images of the same dimensions. There are several methods that generate texture stacks for any material using, for example, smartphones [<xref ref-type="bibr" rid="CIT066">Des+19</xref>; <xref ref-type="bibr" rid="CIT067">DDB20</xref>; <xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT360">Shi+20</xref>; <xref ref-type="bibr" rid="CIT400">VPS21</xref>; <xref ref-type="bibr" rid="CIT337">Rod+23a</xref>], however, these stacks are not tileable by default. Having this data as input, we propose a two-step automatic pipeline to generate tileable texture stacks from arbitrary material input exemplars.</p>
<p>In the first step, described in <xref ref-type="sec" rid="c3-s4">Section 3.4</xref>, we use crops <italic>t</italic> of the input texture <inline-formula><mml:math id="M129" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> to train a Generative Adversarial Network (GAN) able to synthesize novel <italic>untileable</italic> texture stacks <inline-formula><mml:math id="M130" display='block'><mml:msup><mml:mi mathvariant="script">T</mml:mi><mml:mi>&#x2032;</mml:mi></mml:msup></mml:math></inline-formula> using adversarial expansion [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>]. This training framework has been shown to provide state-of-the-art results on single-sample texture synthesis, surpassing previous approaches [<xref ref-type="bibr" rid="CIT394">Uly+16</xref>; <xref ref-type="bibr" rid="CIT028">BJV17</xref>; <xref ref-type="bibr" rid="CIT148">Hen+21</xref>]. Thanks to using a GAN, we learn an implicit representation of the texture parameterized in two neural modules: <inline-formula><mml:math id="M131" display='block'><mml:mi mathvariant="script">G</mml:mi><mml:mtext> and </mml:mtext><mml:mi mathvariant="script">D</mml:mi><mml:mo>.</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msup><mml:mi mathvariant="script">T</mml:mi><mml:mi>&#x2032;</mml:mi></mml:msup></mml:math></inline-formula> is a generator that outputs new untileable textures which double the spatial resolution of the input. <inline-formula><mml:math id="M132" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> is a discriminator that, thanks to the adversarial framework used for training, is able to distinguish real from fake textures.</p>
<p>In the second step, described in <xref ref-type="sec" rid="c3-s5">Section 3.5</xref>, we produce tileable stacks by means of two key ideas: first, we tile a latent space x<sub>0</sub> of the trained generator <inline-formula><mml:math id="M133" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula>, obtaining a novel texture stack showing continuity at the seams intersection <inline-formula><mml:math id="M134" display='block'><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula>. Second, we implement a sampling process using the trained discriminator <inline-formula><mml:math id="M135" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> used as quality metric <inline-formula><mml:math id="M136" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula> to find an optimally tileable texture stack <inline-formula><mml:math id="M137" display='block'><mml:msup><mml:mi mathvariant="script">T</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:math></inline-formula>. An overview of our sampling step is shown in <xref ref-type="fig" rid="fig-3.3">Figure 3.3</xref>.</p>
<fig id="fig-3.3">
<label>Figure 3.3:</label>
<caption><title>Overview of <italic>SeamlessGAN</italic>. A crop <italic>t</italic> of the input texture stack <inline-formula><mml:math id="M119" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula>, is fed to an encoder <inline-formula><mml:math id="M120" display='block'><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula>, which transforms it into a latent space <inline-formula><mml:math id="M121" display='block'><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mi mathvariant="script">E</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula>. We tile this latent space vertically and horizontally, obtaining a latent field F<sub>0</sub>. F<sub>0</sub> is further processed by several residual blocks <inline-formula><mml:math id="M122" display='block'><mml:msub><mml:mi mathvariant="script">R</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo>{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>}</mml:mo><mml:mo>.</mml:mo></mml:math></inline-formula>. The resulting latent variables are transformed by two different decoders, G<sub>A</sub> and G<sub>N</sub>, which output 4 copies of a candidate tileable texture <inline-formula><mml:math id="M123" display='block'><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula>. By cropping the center part of this texture, we obtain a single of those copies, with seamless borders, <inline-formula><mml:math id="M124" display='block'><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula>. Additionally, this texture can be analyzed by a discriminator <inline-formula><mml:math id="M125" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> which provides local estimations of the quality of the synthesis. We introduce a tileability evaluation function <inline-formula><mml:math id="M126" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula>, which, by analyzing two vertical and horizontal search areas <inline-formula><mml:math id="M127" display='block'><mml:msub><mml:mi>s</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math></inline-formula>, is able to detect artifacts that may arise when tiling the texture. This gives us an estimation of how <italic>tileable</italic> the texture is. This estimation can then be used by a sampling algorithm for generating high-quality tileable textures.</title></caption>
<alt-text>Overview of SeamlessGAN. A crop t of the input texture stack T, is fed to an encoder E, which transforms it into a latent space x0&#x2190;E(t). We tile this latent space vertically and horizontally, obtaining a latent field F0. F0 is further processed by several residual blocks Ri,i&#x2208;{1,l}.. The resulting latent variables are transformed by two different decoders, GA and GN, which output 4 copies of a candidate tileable texture T^. By cropping the center part of this texture, we obtain a single of those copies, with seamless borders, T^c. Additionally, this texture can be analyzed by a discriminator D which provides local estimations of the quality of the synthesis. We introduce a tileability evaluation function Q, which, by analyzing two vertical and horizontal search areas sv,sh, is able to detect artifacts that may arise when tiling the texture. This gives us an estimation of how tileable the texture is. This estimation can then be used by a sampling algorithm for generating high-quality tileable textures.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.3.jpg"/>
</fig>
<p>In the following, we first describe our GAN architecture, including our proposal to deal with textures stacks. Then, we describe the sampling process to generate tileable ones.</p>
</sec>
<sec id="c3-s4">
<label>3.4</label>
<title>Self-Supervised Texture Synthesis with Adversarial Expansion</title>
<p>For each input texture <inline-formula><mml:math id="M138" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> we train a GAN, whose generator <inline-formula><mml:math id="M139" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> is able to synthesize novel examples <inline-formula><mml:math id="M140" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> <sup>&#x2032;</sup>. Unlike other GAN frameworks, which take as input a random vector, we use crops <italic>t</italic> of the original stack as input to guide the generative sampling, such that, <inline-formula><mml:math id="M141" display='block'><mml:msup><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mtext> for </mml:mtext><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>T</mml:mi></mml:math></inline-formula>. The training strategy builds upon the work of Zhou <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>], which uses <italic>adversarial expansion</italic> to train the network as follows: First, a target crop <italic>t</italic> &#x2208; <italic>T</italic> of 2<italic>k</italic> &#x00D7; 2<italic>k</italic> pixels is selected from the input stack. Then, from that target crop <italic>t</italic>, a source random crop <italic>t</italic><sub>s</sub> &#x2208; <italic>t</italic> is chosen with a resolution of <italic>k</italic> &#x00D7; <italic>k</italic> pixels. The goal of the generative network will be to synthesize <italic>t</italic> given <italic>t</italic><sub>s</sub>. This learning approach is fully self-supervised. The generative model is trained alongside a discriminator <inline-formula><mml:math id="M145" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula>, which learns to predict whether its inputs are the target crops <italic>t</italic> &#x2208; <italic>T</italic> or the generated samples <inline-formula><mml:math id="M147" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> <sup>&#x2032;</sup> = <inline-formula><mml:math id="M148" display='block'><mml:mrow><mml:mover><mml:mi>T</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. <xref ref-type="fig" rid="fig-3.4">Figure 3.4</xref> shows an overview of the training strategy.</p>
<fig id="fig-3.4">
<label>Figure 3.4:</label>
<caption><title>An overview of our training framework for learning to synthesize texture stacks through adversarial expansion. At each iteration, from an input stack <inline-formula><mml:math id="M151" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> , we randomly crop a target crop <italic>t</italic>, from which we select a source crop <italic>t</italic><sub>s</sub>. The goal of the generator <inline-formula><mml:math id="M152" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> is to estimate <italic>t</italic> given <italic>t<sub>s</sub></italic>. We measure the difference between target and estimated crops using a combination of adversarial, pixel-wise, and perceptual loss functions.</title></caption>
<alt-text>An overview of our training framework for learning to synthesize texture stacks through adversarial expansion. At each iteration, from an input stack T , we randomly crop a target crop t, from which we select a source crop ts. The goal of the generator G is to estimate t given ts. We measure the difference between target and estimated crops using a combination of adversarial, pixel-wise, and perceptual loss functions.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.4.jpg"/>
</fig>
<sec id="c3-s4.s1">
<label><target target-type="page" id="pges_76"/>3.4.1</label>
<title>Network Architecture</title>
<p>SeamlessGAN is comprised of an encoder-decoder convolutional generator <inline-formula><mml:math id="M149" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> with residual connections, and a convolutional discriminator <inline-formula><mml:math id="M150" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula>. So as to be able to synthesize textures of multiple different sizes, the networks are designed to be fully-convolutional. We follow the residual architectural design in [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>], with two extensions: First, building on recent advances in style transfer algorithms, we use <italic>Instance Normalization</italic> [<xref ref-type="bibr" rid="CIT396">UVL16</xref>] before each ReLU non-linearity in the generator and each Leaky-ReLU [<xref ref-type="bibr" rid="CIT260">MHN13</xref>] operation in the discriminator. This allows to use normalization for training the networks without the typical artifacts caused by Batch Normalization [<xref ref-type="bibr" rid="CIT051">Cho+18</xref>; Des+18; NK18; Kur+19]. Second, in order to allow for the synthesis of multiple texture maps at the same time, we propose a variation in the generator architecture. In the following, we describe the details of each component, with a particular focus on the elements that are different from [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>].</p>
<sec id="c3-s4.s1.1">
<title>Generator</title>
<p>The generative architecture proposed is comprised of three main components: An <italic>encoder</italic> <inline-formula><mml:math id="M153" display='block'><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula>, which compresses the information of the input texture <italic>t</italic> into a latent space, <inline-formula><mml:math id="M154" display='block'><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mi mathvariant="script">E</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">t</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula>, with half the spatial resolution of the input texture. Then, <target target-type="page" id="pges_77"/>a set of <italic>residual blocks</italic>, <inline-formula><mml:math id="M155" display='block'><mml:msub><mml:mi mathvariant="script">R</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo>{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">l</mml:mi><mml:mo>}</mml:mo><mml:mo>,</mml:mo></mml:math></inline-formula>, which learn a compact representation of the input texture <inline-formula><mml:math id="M156" display='block'><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mo stretchy="false">&#x2190;</mml:mo><mml:msub><mml:mi mathvariant="script">R</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mrow><mml:mi mathvariant="normal">i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mrow><mml:mi mathvariant="normal">i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. Finally, a stack of <italic>decoders</italic> <bold>G</bold> that transforms the output of the last residual block Rl into an output texture <inline-formula><mml:math id="M158" display='block'><mml:msup><mml:mi mathvariant="script">T</mml:mi><mml:mi>&#x2032;</mml:mi></mml:msup><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mi mathvariant="bold">G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. Residual learning allows for training deeper models with higher levels of visual abstraction and which generate sharper images [<xref ref-type="bibr" rid="CIT139">He+16</xref>; <xref ref-type="bibr" rid="CIT344">RFB15</xref>; <xref ref-type="bibr" rid="CIT176">Iso+17</xref>]. Similar residual generators have been used in unsupervised image-to-image translation problems [<xref ref-type="bibr" rid="CIT460">Zhu+17a</xref>].</p>
<p>Synthesizing a texture stack of multiple maps with a single generative model poses extra challenges over the single map case. Each map represents different properties of the surface, such as geometry or color, resulting in visually and statistically different images. This suggests that independent generative models per map could be needed. However, the texture maps must share pixel-wise coherence, which is not achievable if multiple generative models are used. Inspired by previous work on intrinsic images [<xref ref-type="bibr" rid="CIT179">Jan+17</xref>; <xref ref-type="bibr" rid="CIT439">YS19</xref>; <xref ref-type="bibr" rid="CIT240">LS18</xref>], we propose the use of a generative model that learns a shared representation of all the texture maps, but has different decoders for each of them. The latent space within the generator will thus encode a high-level representation of the texture, which is then decoded in different ways for each different map in the texture stack. Specifically, we train a stack of decoders, <inline-formula><mml:math id="M159" display='block'><mml:mi mathvariant="bold">G</mml:mi><mml:mo mathvariant="bold">=</mml:mo><mml:mrow><mml:mo mathvariant="bold">{</mml:mo><mml:msub><mml:mi mathvariant="bold">G</mml:mi><mml:mi mathvariant="bold">A</mml:mi></mml:msub><mml:mo mathvariant="bold">,</mml:mo><mml:msub><mml:mi mathvariant="bold">G</mml:mi><mml:mi mathvariant="bold">N</mml:mi></mml:msub><mml:mo mathvariant="bold">}</mml:mo></mml:mrow></mml:math></inline-formula>, one for each map of the stack (albedo and normals in our case).</p>
</sec>
<sec id="c3-s4.s1.2">
<title>Discriminator</title>
<p>For the discriminator <inline-formula><mml:math id="M160" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula>, we use a <italic>PatchGAN</italic> architecture [<xref ref-type="bibr" rid="CIT460">Zhu+17a</xref>; <xref ref-type="bibr" rid="CIT176">Iso+17</xref>; <xref ref-type="bibr" rid="CIT230">Led+17</xref>; <xref ref-type="bibr" rid="CIT458">Zho+18</xref>], which, instead of providing a single estimation of the probability of the whole image being real, it classifies this probability for small patches of it. This architecture has several advantages for our problem. First, it provides local assessments of the quality of the synthesized textures, which we exploit for obtaining high-quality textures. Second, its architecture design allows to provide some control over what kind of features are learned: by adding more layers to <inline-formula><mml:math id="M161" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula>, the generated textures are typically of a higher semantic quality but local details may be lost. A comprehensive study on the impact of the depth of <inline-formula><mml:math id="M162" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> can be found in [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>].</p>
</sec>
<sec id="c3-s4.s1.3">
<title>Loss Function</title>
<p>We train the networks following a standard GAN framework [<xref ref-type="bibr" rid="CIT117">Goo+14b</xref>]. We iterate between training <inline-formula><mml:math id="M163" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> with a real sample texture stack <inline-formula><mml:math id="M164" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> and a generated sample <inline-formula><mml:math id="M165" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> <sup>&#x2032;</sup>. The adversarial loss <inline-formula><mml:math id="M166" display='block'><mml:msub><mml:mi mathvariant="bold-script">L</mml:mi><mml:mi mathvariant="bold">adv</mml:mi></mml:msub></mml:math></inline-formula> is extended with three extra loss terms: <inline-formula><mml:math id="M167" display='block'><mml:msubsup><mml:mi mathvariant="bold-script">L</mml:mi><mml:mn mathvariant="bold">1</mml:mn><mml:mi mathvariant="bold">a</mml:mi></mml:msubsup><mml:mo mathvariant="bold">,</mml:mo><mml:msubsup><mml:mi mathvariant="bold-script">L</mml:mi><mml:mn mathvariant="bold">1</mml:mn><mml:mi mathvariant="bold">n</mml:mi></mml:msubsup><mml:mo mathvariant="bold">,</mml:mo></mml:math></inline-formula>, and <inline-formula><mml:math id="M168" display='block'><mml:msub><mml:mi mathvariant="bold-script">L</mml:mi><mml:mtext mathvariant="bold">style</mml:mtext></mml:msub></mml:math></inline-formula>; corresponding respectively to the pixel-wise distance between the generated and target albedo, normals, and a perceptual loss. The perceptual loss <inline-formula><mml:math id="M169" display='block'><mml:msub><mml:mi mathvariant="bold-script">L</mml:mi><mml:mtext mathvariant="bold">style</mml:mtext></mml:msub></mml:math></inline-formula> is computed as the sum of the perceptual losses between target and generated albedo maps, and target and generated normal maps. We follow the <italic>Gram Loss</italic> described in [<xref ref-type="bibr" rid="CIT111">GEB16b</xref>] as our perceptual loss. We weight the total <target target-type="page" id="pges_78"/>style loss by weighting different layers in the same way described in [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>; <xref ref-type="bibr" rid="CIT458">Zho+18</xref>; <xref ref-type="bibr" rid="CIT111">GEB16b</xref>]. Our global loss function thus is defined as:</p>
<p><disp-formula id="Eq003-1"><label>(3.1)</label> <mml:math id="M170" display='block'><mml:mi mathvariant="bold-script">L</mml:mi><mml:mo mathvariant="bold">=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03BB;</mml:mi><mml:mtext mathvariant="bold">adv</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="bold-script">L</mml:mi><mml:mtext mathvariant="bold">adv</mml:mtext></mml:msub><mml:mo mathvariant="bold">+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03BB;</mml:mi><mml:msubsup><mml:mi mathvariant="bold">L</mml:mi><mml:mn mathvariant="bold">1</mml:mn><mml:mi mathvariant="bold">a</mml:mi></mml:msubsup></mml:msub><mml:msubsup><mml:mi mathvariant="bold-script">L</mml:mi><mml:mn mathvariant="bold">1</mml:mn><mml:mi mathvariant="bold">a</mml:mi></mml:msubsup><mml:mo mathvariant="bold">+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03BB;</mml:mi><mml:msubsup><mml:mi mathvariant="bold">L</mml:mi><mml:mn mathvariant="bold">1</mml:mn><mml:mi mathvariant="bold">n</mml:mi></mml:msubsup></mml:msub><mml:msubsup><mml:mi mathvariant="bold-script">L</mml:mi><mml:mn mathvariant="bold">1</mml:mn><mml:mi mathvariant="bold">n</mml:mi></mml:msubsup><mml:mo mathvariant="bold">+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03BB;</mml:mi><mml:mtext mathvariant="bold">style</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="bold-script">L</mml:mi><mml:mtext mathvariant="bold">style</mml:mtext></mml:msub></mml:math></disp-formula></p>
</sec>
</sec>
</sec>
<sec id="c3-s5">
<label>3.5</label>
<title>Tileable Texture Stack Sampling</title>
<sec id="c3-s5.s1">
<label>3.5.1</label>
<title>Latent Space Tiling</title>
<p>After training, the generator <inline-formula><mml:math id="M171" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> is able to synthesize novel samples of the texture given a small exemplar of it. Although these novel samples duplicate the size of the input, they are not tileable by default. Previous work [<xref ref-type="bibr" rid="CIT095">FAW19</xref>] showed that by spatially concatenating different latent spaces in a ProGAN [<xref ref-type="bibr" rid="CIT193">Kar+18</xref>] generator, it is possible to generate textures that contain the visual information of those tiles while seamlessly transitioning between the generated tiles. Inspired by this idea, we spatially repeat (horizontally and vertically) the first latent space x<sub>0</sub> within the generative model, obtaining a latent field F<sub>0</sub>. This field is passed through the residual layers <inline-formula><mml:math id="M172" display='block'><mml:msub><mml:mi mathvariant="bold-script">R</mml:mi><mml:mi mathvariant="bold">l</mml:mi></mml:msub></mml:math></inline-formula> and the decoders <bold>G</bold> to get a texture stack, <inline-formula><mml:math id="M173" display='block'><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula>, that contains four copies of the same texture with seamless transitions between them (see <xref ref-type="fig" rid="fig-3.3">Figure 3.3</xref>). At first, the latent field F<sub>0</sub> shows strong discontinuities at the borders between the four copies. Later, as the texture is passed through the network, these artifacts are progressively transformed into seamless borders by the rest of the residual blocks and the decoder of the model. The intuition behind this idea is that, after training the network with the target texture, each of these latent spaces encode low resolution versions of the input in the spatial domain and semantically-rich information in the deeper layers.</p>
<p>A seamlessly tileable texture stack <inline-formula><mml:math id="M180" display='block'><mml:msub><mml:mover><mml:mi mathvariant="bold">T</mml:mi><mml:mo stretchy="false" mathvariant="bold">^</mml:mo></mml:mover><mml:mi mathvariant="bold">c</mml:mi></mml:msub></mml:math></inline-formula> is obtained by cropping the central region (with an area of 50% of <inline-formula><mml:math id="M181" display='block'><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula>) . The predicted texture <inline-formula><mml:math id="M182" display='block'><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> has 4 &#x00D7; 4 the <target target-type="page" id="pges_79"/>resolution of the input crop <italic>t</italic>. Thus, after the cropping operation, <inline-formula><mml:math id="M183" display='block'><mml:msub><mml:mover><mml:mi mathvariant="bold">T</mml:mi><mml:mo stretchy="false" mathvariant="bold">^</mml:mo></mml:mover><mml:mi mathvariant="bold">c</mml:mi></mml:msub></mml:math></inline-formula> has twice the resolution of the input <italic>t</italic>.</p>
<p>A key parameter to select is the level <inline-formula><mml:math id="M184" display='block'><mml:mi mathvariant="bold-italic">l</mml:mi></mml:math></inline-formula> at which to split the generative process. As shown in <xref ref-type="fig" rid="fig-3.5">Figure 3.5</xref>, generating the latent field F by tiling earlier levels of the latent space forces the network to transform F more times, thus resulting in more seamlessly tileable textures. We confirm the results found in [<xref ref-type="bibr" rid="CIT095">FAW19</xref>] and noticed that generating this latent field at earlier levels of the latent space <inline-formula><mml:math id="M185" display='block'><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo>{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>}</mml:mo><mml:mo>)</mml:mo></mml:math></inline-formula> yields the best visual results, by allowing for smoother and more semantically coherent transitions between tiles. We thus tile the output of the encoder <inline-formula><mml:math id="M186" display='block'><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="script">E</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula>, before transforming it by the residual blocks <inline-formula><mml:math id="M187" display='block'><mml:msub><mml:mi mathvariant="script">R</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>.</p>
<fig id="fig-3.5">
<label>Figure 3.5:</label>
<caption><title>Impact of the latent field level <inline-formula><mml:math id="M174" display='block'><mml:mi mathvariant="bold-italic">l</mml:mi></mml:math></inline-formula> on the quality of the output textures <inline-formula><mml:math id="M175" display='block'><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> = <inline-formula><mml:math id="M176" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula>(<italic>t</italic>), using the same input <italic>t</italic>. Generating the latent field F<sub>&#x03B8;</sub> early layers (<inline-formula><mml:math id="M177" display='block'><mml:mi mathvariant="bold-italic">l</mml:mi></mml:math></inline-formula> = 0) generates the best visual results, later layers either create small artifacts (<inline-formula><mml:math id="M178" display='block'><mml:mi mathvariant="bold-italic">l</mml:mi></mml:math></inline-formula> = 3) or generate unrealistic textures (<inline-formula><mml:math id="M179" display='block'><mml:mi mathvariant="bold-italic">l</mml:mi></mml:math></inline-formula> = 5). A zoom of the central crop is shown as an inset.</title></caption>
<alt-text>Impact of the latent field level l on the quality of the output textures T^ = G(t), using the same input t. Generating the latent field F&#x03B8; early layers (l = 0) generates the best visual results, later layers either create small artifacts (l = 3) or generate unrealistic textures (l = 5). A zoom of the central crop is shown as an inset.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.5.jpg"/>
</fig>
</sec>
<sec id="c3-s5.s2">
<label>3.5.2</label>
<title>Discriminator-guided Sampling</title>
<p>Our strategy to tile the latent space guarantees that the generated texture is a continuous function with smooth transitions between the boundaries of the tiles. However, in contrast to the algorithm in [<xref ref-type="bibr" rid="CIT095">FAW19</xref>], where the latent spaces are drawn from random vectors, ours are encoded representations of input textures. Thus, the selection of the input that the network receives plays an important role on the quality of the output textures. As shown in <xref ref-type="fig" rid="fig-3.6">Figure 3.6</xref>, not all the generated textures are equally valid. This selection can be posed as an optimization problem: <inline-formula><mml:math id="M193" display='block'><mml:msup><mml:mi>t</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>argmax</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mi mathvariant="script">Q</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:math></inline-formula>, where the goal is to find the crop <target target-type="page" id="pges_80"/><italic>t</italic><sup>*</sup> that maximizes the quality <inline-formula><mml:math id="M194" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula> of the generated texture <inline-formula><mml:math id="M195" display='block'><mml:mi mathvariant="script">G</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula>. To solve this optimization problem, one option could be to pose it as an optimization in the generative latent space. However, we found that a simpler solution based on sampling already provides satisfactory results. The remainder of this section explains the sampling process and the quality metric.</p>
<fig id="fig-3.6">
<label>Figure 3.6:</label>
<caption><title>Outputs of <inline-formula><mml:math id="M188" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> for textures of different qualities. (a) The generated albedos <inline-formula><mml:math id="M189" display='block'><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> are artifact-free on the search areas (b), thus the discriminator <inline-formula><mml:math id="M190" display='block'><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mn>1</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is fooled to believe <inline-formula><mml:math id="M191" display='block'><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> cannot fool the discriminator, which finds artifacts on the central areas <inline-formula><mml:math id="M192" display='block'><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> (d).</title></caption>
<alt-text>Outputs of D for textures of different qualities. (a) The generated albedos T^1 are artifact-free on the search areas (b), thus the discriminator D(T^1) is fooled to believe T^2 cannot fool the discriminator, which finds artifacts on the central areas D(T^2) (d).</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.6.jpg"/>
</fig>
<p><bold>Sampling</bold> By using a fully-convolutional GAN, our model can generate textures of any size, hardware being the only limiting factor. This is key for tileable texture synthesis as, even if a given texture is seamlessly tileable, larger textures require fewer repetitions to cover the same spatial area, which ultimately results in fewer repeating artifacts, as illustrated in <xref ref-type="fig" rid="fig-3.7">Figure 3.7</xref> (a). There are two main challenges when finding tileable textures: the input texture needs to contain the distribution of the existing repeating patterns, and the input tile itself must not create strong artifacts when tiling the latent spaces of the generator.</p>
<fig id="fig-3.7">
<label>Figure 3.7:</label>
<caption><title><italic>SeamlessGAN</italic> can generate multiple tileable outputs from the same sample <inline-formula><mml:math id="M196" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula>. By cropping different parts of <inline-formula><mml:math id="M197" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> , we can feed the generator <inline-formula><mml:math id="M198" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> with different inputs <italic>t</italic><sub>c</sub> (shown on the left), generating tileable textures <inline-formula><mml:math id="M199" display='block'><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula> of different sizes. On the example at the top (a), the synthesized images are tiled so they cover the same spatial extent, which shows that, even if the textures are tileable, larger textures generate fewer repeating artifacts. (b) Shows the tile resulting from sampling different parts of the input image.</title></caption>
<alt-text>SeamlessGAN can generate multiple tileable outputs from the same sample T. By cropping different parts of T , we can feed the generator G with different inputs tc (shown on the left), generating tileable textures T^c of different sizes. On the example at the top (a), the synthesized images are tiled so they cover the same spatial extent, which shows that, even if the textures are tileable, larger textures generate fewer repeating artifacts. (b) Shows the tile resulting from sampling different parts of the input image.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.7.jpg"/>
</fig>
<p><target target-type="page" id="pges_81"/>Our goal is thus to find the largest possible tileable texture stack. To do so, we sample multiple candidate crops for a given crop size <inline-formula><mml:math id="M200" display='block'><mml:mi>c</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mo>min</mml:mo></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mo>max</mml:mo></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> is the resolution of the input <inline-formula><mml:math id="M201" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula>. We sample crop sizes starting at the largest possible size <inline-formula><mml:math id="M202" display='block'><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mo>max</mml:mo></mml:msub></mml:math></inline-formula>, and stop when we find a suitable candidate according to the tileability metric <inline-formula><mml:math id="M203" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula>. As shown in <xref ref-type="fig" rid="fig-3.7">Figure 3.7</xref> (b), this sampling mechanism also allows us to generate multiple tileable candidates for a single exemplar if we choose different parts of <inline-formula><mml:math id="M204" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> as input.</p>
<p><bold>Discriminator-guided Quality Function,</bold> <inline-formula><mml:math id="M205" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula> The second component of our sampling strategy is the quality function used to determine whether a stack is tileable or not. We observed that the artifacts appear on vertical and horizontal frames around the center of the textures (<xref ref-type="fig" rid="fig-3.6">Figure 3.6</xref>). This is likely caused by strong discontinuities or gradients on the same areas of the tiled latent spaces, which the rest of the generative network fails to transform into realistic textures. Following recent work on generative models [<xref ref-type="bibr" rid="CIT354">SSK20</xref>], we use the discriminator <inline-formula><mml:math id="M206" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> as a semantic-aware error metric that can be exploited for detecting local artifacts in the generated textures. This can be done in our case because the global loss function contains pixel-wise, <italic>style</italic>, and adversarial losses. The adversarial loss learns the semantics of the texture, whereas the <inline-formula><mml:math id="M207" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and style losses model color distances or repeated patterns [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>]. This combination of loss functions allows us to balance textural, semantic, and perceptual color properties.</p>
<p>We thus design a quality evaluation function <inline-formula><mml:math id="M208" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula> that estimates if the generated texture stack <inline-formula><mml:math id="M209" display='block'><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula> is <italic>tileable</italic>, looking for artifacts on a central area, <inline-formula><mml:math id="M210" display='block'><mml:mi>S</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of the discriminator. This area is composed of two regions, <inline-formula><mml:math id="M211" display='block'><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>&#x222A;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math></inline-formula>, where <inline-formula><mml:math id="M212" display='block'><mml:msub><mml:mi>S</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> is a vertical area, and <inline-formula><mml:math id="M213" display='block'><mml:msub><mml:mi>S</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math></inline-formula> is an horizontal one, both centered on the output of the discriminator (<xref ref-type="fig" rid="fig-3.3">Figure 3.3</xref>). The function <inline-formula><mml:math id="M214" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula> leverages the fact that <inline-formula><mml:math id="M215" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> outputs 0 when it believes a patch to be synthetic. However, as the values are sample-dependent, we establish a threshold <inline-formula><mml:math id="M216" display='block'><mml:mi>&#x03C4;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="normal">&#x211D;</mml:mi></mml:math></inline-formula> using the values of the rest of the image <italic>S</italic> as a reference <inline-formula><mml:math id="M217" display='block'><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mo>min</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M218" display='block'><mml:msub><mml:mi>S</mml:mi><mml:mi>r</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2229;</mml:mo><mml:mover><mml:mi>S</mml:mi><mml:mo stretchy="false">&#xAF;</mml:mo></mml:mover></mml:math></inline-formula> is the remaining part of the image, and <inline-formula><mml:math id="M219" display='block'><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="normal">&#x211D;</mml:mi></mml:math></inline-formula> is a threshold that allows to control the sensitiveness of <inline-formula><mml:math id="M220" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula>. Consequently, <inline-formula><mml:math id="M221" display='block'><mml:mi mathvariant="script">Q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover><mml:mi mathvariant="script">T</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is 1 if min(<inline-formula><mml:math id="M222" display='block'><mml:msub><mml:mi>s</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula>) and min(<inline-formula><mml:math id="M223" display='block'><mml:msub><mml:mi>s</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math></inline-formula>) are greater or equal than <inline-formula><mml:math id="M224" display='block'><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>, considering the texture as <italic>tileable</italic>, and 0 otherwise. The goal of <inline-formula><mml:math id="M225" display='block'><mml:mi mathvariant="script">Q</mml:mi></mml:math></inline-formula> is to estimate whether or not the central areas, where discontinuities may be present, are as realistic as the rest of the texture. It works as a classification function which returns a high value if the texture is identified as real by the discriminator in such areas. This allows us to create a sampling strategy that distinguishes between images with local artifacts from seamless textures. By using minimum values instead of the mean estimation of the discriminator, our quality metric focuses on detecting artifacts which may arise during the latent space tiling operation.</p>
</sec>
</sec>
<sec id="c3-s6">
<label>3.6</label>
<title>Implementation Details</title>
<p>As described in <xref ref-type="sec" rid="c3-s4">Section 3.4</xref>, we follow a standard GAN training framework, by iterating between training the discriminator and the generator. We use a batch <target target-type="page" id="pges_82"/>size of 1, and an input size of <italic>k</italic> = 128. All weights are initialized by sampling a Gaussian distribution <inline-formula><mml:math id="M226" display='block'><mml:mi mathvariant="script">N</mml:mi></mml:math></inline-formula> (0, 0.02), following standard practice [<xref ref-type="bibr" rid="CIT460">Zhu+17a</xref>]. <inline-formula><mml:math id="M227" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> has <inline-formula><mml:math id="M228" display='block'><mml:mi mathvariant="bold-italic">l</mml:mi></mml:math></inline-formula> = 5 residual blocks with ReLU activations, and <inline-formula><mml:math id="M229" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula> is comprised of 5 convolutional layers with Leaky-ReLU non-linearities, and a Sigmoid operation at the end. We use a stride of 2 for the downsampling operations in <inline-formula><mml:math id="M230" display='block'><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula> and transposed convolutions [<xref ref-type="bibr" rid="CIT254">LSD15</xref>] for upsampling in the decoders <bold>G</bold>. We weight each part of the loss function as: <inline-formula><mml:math id="M231" display='block'><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>and</mml:mtext><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:msubsup><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn><mml:mi>a</mml:mi></mml:msubsup></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:msubsup><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:msubsup></mml:msub><mml:mo>=</mml:mo><mml:mn>10</mml:mn><mml:mo>.</mml:mo></mml:math></inline-formula>.</p>
<p>The networks are trained for 50000 iterations using Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>], with an initial learning rate of 0.0002, which is divided by 5, after iterations 30000 and 40000. Aside from random cropping, we do not use any other type of data augmentation method. The models are trained and evaluated using a single NVIDIA GeForce GTX 1080Ti. Even if training takes around 40 minutes for each texture stack, once trained, the generator can generate individual new samples in milliseconds, before checking for tileability. We use PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>] as our learning framework. To accelerate the training process, we leverage mixed precision training and automatic gradient scaling [<xref ref-type="bibr" rid="CIT274">Mic+18</xref>]. Every operation in the training pipeline is done natively in GPU using Torchvision [<xref ref-type="bibr" rid="CIT262">MR10</xref>]. These optimizations allow us to train the networks one order of magnitude faster than previous methods [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>]. The input textures <inline-formula><mml:math id="M232" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> tested in this chapter have, on average, 500 pixels in their larger dimension, for which our method can generate tiles of at most 1000 pixels. We use Photometric Stereo [<xref ref-type="bibr" rid="CIT174">Ike81</xref>] for computing the normals of fabric textures, and artist-generated normals for the other samples.</p>
<p>We tile the latent space at its earliest level <inline-formula><mml:math id="M233" display='block'><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">F</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:math></inline-formula>, as it provides the best quality results, as shown in <xref ref-type="sec" rid="c3-s5.s1">Section 3.5.1</xref>. For identifying the tileability of textures, we use a search area <italic>S</italic> that spans a 20% of each spatial dimension of the textures. For all the results shown, &#x03B3; = 1. We choose <inline-formula><mml:math id="M234" display='block'><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mo movablelimits="true">min</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>100</mml:mn><mml:mtext> pixels </mml:mtext></mml:math></inline-formula>, and <inline-formula><mml:math id="M235" display='block'><mml:msub><mml:mi>c</mml:mi><mml:mo>max</mml:mo></mml:msub></mml:math></inline-formula> equal to the resolution of the whole input texture. This means that for the level of <inline-formula><mml:math id="M236" display='block'><mml:msub><mml:mi>c</mml:mi><mml:mo>max</mml:mo></mml:msub></mml:math></inline-formula>, only one texture is sampled. For each <inline-formula><mml:math id="M237" display='block'><mml:mi>c</mml:mi><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mo>max</mml:mo></mml:msub></mml:math></inline-formula>, we sample 3 different random crops. Typically, for the results shown in this chapter, a tileable texture stack is found at high-resolution crop sizes, <inline-formula><mml:math id="M238" display='block'><mml:mi>c</mml:mi><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mo>max</mml:mo></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula>, thus needing to sample and evaluate less than thirty random crops before finding a satisfactory solution. This entire sampling procedure takes less than five minutes.</p>
</sec>
<sec id="c3-s7">
<label><target target-type="page" id="pges_83"/>3.7</label>
<title>Experiments</title>
<p>In this section, we evaluate two of our main contributions: first, in <xref ref-type="sec" rid="c3-s7.s1">Section 3.7.1</xref> we study the impact of the design choices of the generator <inline-formula><mml:math id="M239" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> to synthesize a multi-layer texture stack, and second, the quality of the tileable texture synthesis through several ablation studies. We study the inter-map consistency in <xref ref-type="sec" rid="c3-s7.s2">Section 3.7.2</xref>, and the impact of the loss function in <xref ref-type="sec" rid="c3-s7.s3">Section 3.7.3</xref>. Finally, we compare our method with methods on tileable texture synthesis (<xref ref-type="sec" rid="c3-s8">Section 3.8</xref>).</p>
<sec id="c3-s7.s1">
<label>3.7.1</label>
<title>Generator <inline-formula><mml:math id="M240" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> Design</title>
<p>One of the main challenges when synthesizing texture stacks using a single generator is to preserve the low-level details of each of the texture maps whilst maintaining the local spatial coherence between them; if this coherence is lost, renders that use the synthesized maps will show artifacts or look unrealistic. We propose two different variations of the single-map generative architecture <inline-formula><mml:math id="M241" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> presented in [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>], each of which makes different assumptions on how the synthesis should be learned taking into account the particular semantics and purpose of each map. A diagram of each proposed architecture is shown in <xref ref-type="fig" rid="fig-3.10">Figure 3.10</xref>. For a fair evaluation, we follow the criteria that both networks must have approximately the same number of trainable parameters.</p>
<p>Our baseline, <inline-formula><mml:math id="M242" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>, treats the texture stack as a multiple-channel input image, and entangles every texture map in the same layers. It assumes that the maps in the stack share most of the structural information and, as such, there is no need to generate them separately. Thus, the last layer in the decoder outputs every texture map. Our proposed alternative architecture, <inline-formula><mml:math id="M243" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> finds a shared representation of each texture map, but has a separate decoder for each of them. As such, the residual blocks are shared for all the texture stack, but each decoder can be optimized for the semantics and statistics of each particular map.</p>
<p>To quantitatively evaluate which architecture produces the highest quality output we compare the original texture with the generated one using standard metrics: SSIM, Si-FID, and LPIPS. The <italic>Structural Similarity Index Measure (SSIM)</italic> [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>] is a perceptual-aware metric, working on the pixel space, that measures the similarity in structural information, and may be appropriate to evaluate synthesized textures. The Si-FID [<xref ref-type="bibr" rid="CIT359">SDM19</xref>] is a single image extension of the <italic>Fr&#x00E9;chet Inception Score (FID)</italic> [<xref ref-type="bibr" rid="CIT153">Heu+17</xref>], which measures the difference in deep latent statistics between natural and artificially generated images. Finally, we use the <italic>Learned Perceptual Image Patch Similarity (LPIPS)</italic> [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] as a perceptual distance metric in deep image spaces. This metric is widely used for evaluating generative models [<xref ref-type="bibr" rid="CIT197">Kar+20b</xref>; <xref ref-type="bibr" rid="CIT169">Hua+18b</xref>; <xref ref-type="bibr" rid="CIT042">Cha+19</xref>; <xref ref-type="bibr" rid="CIT008">Alm+18</xref>; <xref ref-type="bibr" rid="CIT263">Mar+20</xref>].</p>
<p><target target-type="page" id="pges_84"/>The baseline <inline-formula><mml:math id="M244" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> shows more artifacts than <inline-formula><mml:math id="M245" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, most likely due to the fact that the generation parts of the network are not fully separated. A quantitative evaluation is shown in <xref ref-type="table" rid="c3-tab1">Table 3.1</xref>, showing that <inline-formula><mml:math id="M246" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> outperforms <inline-formula><mml:math id="M247" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> from perceptual and statistical standpoints. Those metrics require the input and target images to have the same spatial dimensions. To obtain this, we crop the 50% center area of each generated stack, which doubles the dimensions of the inputs due to the adversarial expansion approach we follow, and compare them with the input textures. In summary, <inline-formula><mml:math id="M248" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> allows to generate textures that better preserve the properties of the input images without any additional computational cost and without requiring to modify the loss function, or the discriminator design.</p>
<table-wrap id="c3-tab1">
<label>Table 3.1:</label>
<caption><title>Quantitative comparison between our variations of the generator <inline-formula><mml:math id="M249" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula>. We show the average results on different metrics across different texture stacks, separated by maps. As shown, <inline-formula><mml:math id="M250" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> outperforms its baseline across metrics and texture maps. <inline-formula><mml:math id="M251" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> only yields better scores at the Si-FID metric on the normal map. Higher is better for SSIM, while lower is better for Si-FID and LPIPS</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center"/>
<th valign="top" align="left"><p>SSIM [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>] &#x2191;</p></th>
<th valign="top" align="left"><p>Si-FID [<xref ref-type="bibr" rid="CIT359">SDM19</xref>] &#x2193;</p></th>
<th valign="top" align="left"><p>LPIPS [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] &#x2193;</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" rowspan="2"><p><inline-formula><mml:math id="M252" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula></p></td>
<td valign="top" align="left"><p>Albedo</p></td>
<td valign="top" align="center"><p>0.2774</p></td>
<td valign="top" align="center"><p>0.3462</p></td>
<td valign="top" align="center"><p>0.4826</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Normals</p></td>
<td valign="top" align="center"><p>0.2364</p></td>
<td valign="top" align="center"><p><bold>0.3636</bold></p></td>
<td valign="top" align="center"><p>0.4870</p></td>
</tr>
<tr>
<td valign="top" align="left" rowspan="2"><p><inline-formula><mml:math id="M253" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula></p></td>
<td valign="top" align="left"><p>Albedo</p></td>
<td valign="top" align="center"><p><bold>0.3123</bold></p></td>
<td valign="top" align="center"><p><bold>0.3275</bold></p></td>
<td valign="top" align="center"><p><bold>0.4377</bold></p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Normals</p></td>
<td valign="top" align="center"><p><bold>0.2921</bold></p></td>
<td valign="top" align="center"><p>0.4019</p></td>
<td valign="top" align="center"><p><bold>0.4340</bold></p></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="c3-s7.s2">
<label>3.7.2</label>
<title>Inter-map Consistency</title>
<p>Whilst our generators <inline-formula><mml:math id="M254" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> can generate high-quality tileable pairs of albedo and normal maps, there is no guarantee that those maps are pixel-wise coherent, as no part in the loss function explicitly accounts for this relationship. The <inline-formula><mml:math id="M255" display='block'><mml:msub><mml:mi mathvariant="bold-script">L</mml:mi><mml:mtext mathvariant="bold">style</mml:mtext></mml:msub></mml:math></inline-formula> loss is computed separately for each map in the texture stack, which may generate non-coherent gradients. Furthermore, our architecture of choice <inline-formula><mml:math id="M256" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> separates the decoding of each map in different parts of its architecture, which may hinder the generation of spatially-coherent maps. Nevertheless, we show that this is not a problem in practice. <xref ref-type="fig" rid="fig-3.8">Figure 3.8</xref> shows crops of our synthesized texture maps. It can be seen that pixel-wise coherency between maps is preserved even for challenging geometric structures. We argue that the role of the discriminator is key to detect inter-layer inconsistencies by yielding lower probabilities for non-coherent maps. Even if separating the decoders may increase the risk of incoherent texture maps, this coherence is forced by the discriminator during training. To test this empirically, for the textures in <xref ref-type="fig" rid="fig-3.8">Figure 3.8</xref>, we fed the discriminator with a crop of the original stack, a crop with the normals translated 5 <target target-type="page" id="pges_85"/>px, and translated 100 px, for which we obtain average values of 0.99, 0.35, and 0.32, respectively. This suggests that, since during training, the discriminator receives the entire texture stack at once, it learns to identify spatial inconsistencies between maps.</p>
<fig id="fig-3.8">
<label>Figure 3.8:</label>
<caption><title>Crops of synthesized textures using our method. We observe that pixel-wise coherence between maps is preserved.</title></caption>
<alt-text>Crops of synthesized textures using our method. We observe that pixel-wise coherence between maps is preserved.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.8.jpg"/>
</fig>
</sec>
<sec id="c3-s7.s3">
<label>3.7.3</label>
<title>Loss Function Ablation Study</title>
<p>As described in <xref ref-type="sec" rid="c3-s4">Section 3.4</xref>, the generator <inline-formula><mml:math id="M257" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> is trained to minimize a loss function, which is comprised of adversarial, perceptual and pixel-wise components. Each of these have different impacts on the output synthesis. We hereby study the impact of each of these components.</p>
<p>To better isolate the impact of each metric, we perform this study using a limited generator that only outputs albedo maps. As we show in <xref ref-type="fig" rid="fig-3.9">Figure 3.9</xref>, the adversarial loss <inline-formula><mml:math id="M259" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> provides high-level semantic consistency, the <inline-formula><mml:math id="M260" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> norm acts as a regularization method which removes artifacts, while the perceptual loss <inline-formula><mml:math id="M261" display='block'><mml:msub><mml:mi mathvariant="bold-script">L</mml:mi><mml:mtext mathvariant="bold">style</mml:mtext></mml:msub></mml:math></inline-formula> adds additional details to fully represent the texture. Similar findings are reported in [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>]. None of these loss functions yield compelling or artifact-free results on isolation, being the adversarial loss the most important factor of the global function.</p>
<fig id="fig-3.9">
<label>Figure 3.9:</label>
<caption><title>Ablation study on the impact of the loss function on the quality of the synthesized textures. We evaluate each network using the same input, marked using a blue box. As shown, training <inline-formula><mml:math id="M258" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> using the full loss function yields the best results. Each component is weighted by a <italic>&#x03BB;</italic>, as specified in <xref ref-type="sec" rid="c3-s6">Section 3.6</xref>.</title></caption>
<alt-text>Ablation study on the impact of the loss function on the quality of the synthesized textures. We evaluate each network using the same input, marked using a blue box. As shown, training G using the full loss function yields the best results. Each component is weighted by a &#x03BB;, as specified in Section 3.6.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.9.jpg"/>
</fig>
</sec>
</sec>
<sec id="c3-s8">
<label><target target-type="page" id="pges_86"/>3.8</label>
<title>Results and Comparisons</title>
<p>Most tileable texture synthesis methods have not shown results in synthesizing texture stacks. While adding this capability is reasonably easy for some <italic>non parametric</italic> methods [<xref ref-type="bibr" rid="CIT241">Li+20</xref>; <xref ref-type="bibr" rid="CIT341">Rod+19</xref>], others require major changes in their design. In particular, as we have shown in <xref ref-type="sec" rid="c3-s7.s1">Section 3.7.1</xref>, expanding the architecture of deep learning methods for generating more than one map requires special attention to balance the map&#x2019;s interdependence with the network&#x2019;s capacity and expected quality. Therefore, in this section, in order to be able to compare our method with other works, we use a reduced version of our generator that only outputs a single albedo map.</p>
<fig id="fig-3.10">
<label>Figure 3.10:</label>
<caption><title>A comparison of the results of our proposed architectures. Separating the decoders of the network for each map (<inline-formula><mml:math id="M262" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>) outperforms the joint-decoder baseline (<inline-formula><mml:math id="M263" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>) for both texture maps in the stacks. It generally allows for more varied albedo maps, with a more correct structure and style; as well as more accurate and sharper normal maps. The outputs in this figure are not tiled, for improved visibility of the artifacts created by <inline-formula><mml:math id="M264" display='block'><mml:msub><mml:mi mathvariant="script">G</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>. In blue, we show an up-sampled crop of the output textures.</title></caption>
<alt-text>A comparison of the results of our proposed architectures. Separating the decoders of the network for each map (G2) outperforms the joint-decoder baseline (G1) for both texture maps in the stacks. It generally allows for more varied albedo maps, with a more correct structure and style; as well as more accurate and sharper normal maps. The outputs in this figure are not tiled, for improved visibility of the artifacts created by G1. In blue, we show an up-sampled crop of the output textures.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.10.jpg"/>
</fig>
<p><target target-type="page" id="pges_87"/><bold>Qualitative Analysis</bold> First, we compare with the Texture Stationarization algorithm proposed in [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>] using examples taken from their own dataset, as their code is not publicly available. A key difference between both methods is that, while their method aims to maximize the stationarity properties of the textures using external metrics, our method learns them from the texture itself, using a self-supervised approach. This results in our method better preserving the content of the original input. As shown in <xref ref-type="fig" rid="fig-3.11">Figure 3.11</xref>, our method shows compelling results for the kind of textures shown in their paper. In the <italic>fence</italic> example, our methods better preserves the vertical straight lines. For the <italic>wall</italic> example, both methods shows compelling results, capturing a different repetition pattern.</p>
<fig id="fig-3.11">
<label>Figure 3.11:</label>
<caption><title>A comparison with the work on Texture Stationarization by Moritz <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>], using textures from their dataset. The results are tiled 2 &#x00D7; 2 times to help visualization.</title></caption>
<alt-text>A comparison with the work on Texture Stationarization by Moritz et al. [Mor+17], using textures from their dataset. The results are tiled 2 &#x00D7; 2 times to help visualization.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.11.jpg"/>
</fig>
<p>As Moritz <italic>et al.</italic>&#x2019;s dataset contains mainly textures of human-made environments, we gather a different set of images of greater variety in their regularity and content. <xref ref-type="fig" rid="fig-3.12">Figure 3.12</xref> shows a comprehensive comparison with the other methods. The work by Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT241">Li+20</xref>], based on <italic>Graph-Cuts</italic>, works reasonably well if the transformation can be done locally at the seams but fails when the required changes are global, as happens in the <italic>basket</italic>. The <italic>Repeated Pattern Detection</italic> algorithm proposed by Rodriguez-Pardo <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>] is not able to handle many of these challenging cases, which do not represent grid-like textures, as in the <italic>bananas</italic> or <italic>stars</italic>. The Periodic Spatial GAN (PSGAN) proposed by Bermann <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT028">BJV17</xref>] generates artifacts, not maintaining the content of the texture, as in the <italic>daisies</italic> or the <italic>hieroglyph</italic> examples. A similar effect is observable in the method of [<xref ref-type="bibr" rid="CIT296">Nik+21</xref>], which uses self-organizing representations through neural cellular automata. The output of the method is seamless but the integrity of the original texture is not preserved in many cases, for example, in the <italic>hieroglyph</italic> or the <italic>stars</italic>.</p>
<fig id="fig-3.12">
<label>Figure 3.12:</label>
<caption><title>Comparison with previous methods. On the left column, we show the input textures. From left to right, the synthesized results of Histogram Blending, by Deloit <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT063">DH19</xref>], the GraphCuts algorithm by Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT241">Li+20</xref>], Repeated Pattern Detection by Rodriguez-Pardo <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>], PSGAN by Bergmann <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT028">BJV17</xref>], Self-Organizing Textures by Niklasson <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT296">Nik+21</xref>]; and ours. Outputs are tiled a similar number of times (at least twice in each dimension) for better visualization. Our method generally captures better the overall structure of the texture, while providing seamless and seman-tically coherent borders, for enhancing tileability. From top to bottom, we sample c = 4, 12, 5, 17, 13, 1, 19 and 43 input crops before obtaining a tileable texture.</title></caption>
<alt-text>Comparison with previous methods. On the left column, we show the input textures. From left to right, the synthesized results of Histogram Blending, by Deloit et al. [DH19], the GraphCuts algorithm by Li et al. [Li+20], Repeated Pattern Detection by Rodriguez-Pardo et al. [Rod+19], PSGAN by Bergmann et al. [BJV17], Self-Organizing Textures by Niklasson et al. [Nik+21]; and ours. Outputs are tiled a similar number of times (at least twice in each dimension) for better visualization. Our method generally captures better the overall structure of the texture, while providing seamless and semantically coherent borders, for enhancing tileability. From top to bottom, we sample c = 4, 12, 5, 17, 13, 1, 19 and 43 input crops before obtaining a tileable texture.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.12.jpg"/>
</fig>
<p>Our method favors keeping high-level and bigger semantic structures of the textures, resulting in larger samples with more variety. The modification applied to the texture is the minimum required to make it seamlessly tileable. It provides high-quality outputs for regular textures with uneven illumination and perspective distortion, such as the <italic>basket</italic>, textures which are greatly irregular, such as the <italic>bananas</italic>, and stochastic ones, like the <italic>strawberry</italic>. While non-parametric synthesis methods are typically computationally cheap, obtaining results in less than one minute [<xref ref-type="bibr" rid="CIT241">Li+20</xref>; Rod+19; DH19], at the cost of quality and generality, parametric methods are more expensive and slower. Our method is competitive in computational cost when compared to other parametric synthesis methods. Our entire training and sampling process takes less than 45 minutes. In comparison, PSGAN [<xref ref-type="bibr" rid="CIT028">BJV17</xref>] needs 4 hours to train, the work by Zhou <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>] requires 6 hours, and the method of Niklasson <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT296">Nik+21</xref>] takes 35 minutes to generate textures limited to 200 &#x00D7; 200 pixels. On the contrary, our method is not limited by the size of the input texture and can synthesize tiles of any size thanks to the fully-convolutional architecture, hardware being the only limiting factor.<target target-type="page" id="pges_88"/></p>
<p><target target-type="page" id="pges_89"/><bold>Quantitative Comparison</bold> Following a similar evaluation scheme as proposed in <xref ref-type="sec" rid="c3-s7.s1">Section 3.7.1</xref> for measuring the differences between input and synthesized images, a quantitative evaluation is shown in <xref ref-type="table" rid="c3-tab2">Table 3.2</xref>. We use Moritz&#x2019;s dataset for a fair comparison with every method. The metrics we use for this experiment require that the input and synthesized images have the same resolution. To achieve this, and, in order to account for artifacts in the borders of the generated textures, we tile the synthesized images until they cover the same resolution as their corresponding inputs.</p>
<table-wrap id="c3-tab2">
<label>Table 3.2:</label>
<caption><title>Quantitative comparison between different methods. We show the average results on different perceptual metrics across a variety of textures.As shown, SeamlessGAN consistently outperforms its counterparts in every studied metric. We tile the outputs until they match the spatial resolution of the input examples. Higher is better for SSIM, while lower is better for Si-FID and LPIPS. We use a color code to highlight <bold>best</bold> and <bold>worst</bold> cases.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center"><p>SSIM [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>] &#x2191;</p></th>
<th valign="top" align="center"><p>Si-FID [<xref ref-type="bibr" rid="CIT359">SDM19</xref>] &#x2193;</p></th>
<th valign="top" align="center"><p>LPIPS [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] &#x2193;</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p>Deloit <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT063">DH19</xref>]</p></td>
<td valign="top" align="center"><p><bold>0.1424</bold></p></td>
<td valign="top" align="center"><p>1.3471</p></td>
<td valign="top" align="center"><p><bold>0.6207</bold></p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT241">Li+20</xref>]</p></td>
<td valign="top" align="center"><p>0.2086</p></td>
<td valign="top" align="center"><p>0.9529</p></td>
<td valign="top" align="center"><p>0.5818</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Moritz <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>]</p></td>
<td valign="top" align="center"><p>0.1968</p></td>
<td valign="top" align="center"><p>0.7620</p></td>
<td valign="top" align="center"><p>0.5171</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Rodriguez-Pardo <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>]</p></td>
<td valign="top" align="center"><p>0.2144</p></td>
<td valign="top" align="center"><p>1.2958</p></td>
<td valign="top" align="center"><p>0.5137</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Bergmann <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT028">BJV17</xref>]</p></td>
<td valign="top" align="center"><p>0.1723</p></td>
<td valign="top" align="center"><p><bold>1.4355</bold></p></td>
<td valign="top" align="center"><p>0.5624</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Niklasson <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT296">Nik+21</xref>]</p></td>
<td valign="top" align="center"><p>0.1753</p></td>
<td valign="top" align="center"><p>1.3171</p></td>
<td valign="top" align="center"><p>0.5328</p></td>
</tr>
<tr>
<td valign="top" align="left"><p><bold>SeamlessGAN</bold></p></td>
<td valign="top" align="center"><p><bold>0.2341</bold></p></td>
<td valign="top" align="center"><p><bold>0.6311</bold></p></td>
<td valign="top" align="center"><p><bold>0.4792</bold></p></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Interestingly, methods based on patches [<xref ref-type="bibr" rid="CIT342">RB19</xref>; <xref ref-type="bibr" rid="CIT281">Mor+17</xref>; <xref ref-type="bibr" rid="CIT241">Li+20</xref>] obtain better SSIM and Si-FID scores than previous deep learning-based methods [<xref ref-type="bibr" rid="CIT028">BJV17</xref>; <xref ref-type="bibr" rid="CIT296">Nik+21</xref>]. This difference is not seen in the LPIPS metric. This suggests that patch-based methods preserve better the structure of the input textures than previous neural parametric models. Our method, by combining a variety of loss functions, allows for better preservation of the style and semantic content of the generated textures. Furthermore, our latent space manipulation algorithm allows for seamless borders between tiles, outperforming both previous non-parametric and parametric methods in perceptual and structural similarity. The magnitude of the Si-FID metric varies significantly between the values obtained in <xref ref-type="table" rid="c3-tab1">Table 3.1</xref> and <xref ref-type="table" rid="c3-tab2">Table 3.2</xref>, which indicates that this metric may be overly sensitive to the global statistics of the dataset.</p>
<p><bold>Limitations and Discussion</bold> Our method is inherently limited by the capabilities of the adversarial expansion technique to learn the implicit structure of the given texture. That is, if the input does not show enough regularity, the adversarial expansion fails, as shown in <xref ref-type="fig" rid="fig-3.13">Figure 3.13</xref>. It is interesting to see that the synthesis of the <italic>dotted</italic> texture fails to reproduce the larger dots, as they are very scarce. However, the remaining structure is very well represented.</p>
<fig id="fig-3.13">
<label>Figure 3.13:</label>
<caption><title>A limitation of our texture synthesis algorithm: Left: Input texture stacks; right: synthesized stacks. The network fails to replicate the pattern if the occurrence is not frequent enough, as we can see in the base color of these two examples. On the other hand, the synthesized normal map is consistent as its repetitive structure is seen frequently by the network.</title></caption>
<alt-text>A limitation of our texture synthesis algorithm: Left: Input texture stacks; right: synthe-sized stacks. The network fails to replicate the pattern if the occurrence is not frequent enough, as we can see in the base color of these two examples. On the other hand, the synthesized normal map is consistent as its repetitive structure is seen frequently by the network.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.13.jpg"/>
</fig>
<p>Tiling single texture maps, even if they contain seamless borders, may generate perceptible repetitions. This is an intrinsic limitation of any single sample tileable texture synthesis. Approaches such as <italic>Wang Tiles</italic> [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>] can tackle these limitations, but have additional disadvantages like increased memory and run-time consumption, rendering them less practical for real-time applications. Alternatively, procedural methods [<xref ref-type="bibr" rid="CIT123">Gue+20</xref>] are the most effective way to generate material samples preserving textural properties at multiple scales; however, present several limitations in the range of materials that can be generated to make them fully usable.</p>
</sec>
<sec id="c3-s9">
<label>3.9</label>
<title>Conclusions and Future Work</title>
<p>In this chapter, we have proposed a deep parametric texture synthesis framework capable of synthesizing textures into tileable single-tiles, by combining recent advances in deep texture synthesis, adversarial neural networks, and latent spaces manipulation. Our results show that our method can <target target-type="page" id="pges_90"/><target target-type="page" id="pges_91"/>generate visually pleasing results for images with different levels of regularity and homogeneity. This work is the first method capable of exploiting properties of deep latent spaces within neural networks for generating seamless textures, and opens the opportunity for end-to-end tileable texture synthesis methods without the need for manual input. Comparisons with previous state-of-the-art methods show that our method provides results which better maintain the semantic properties of the textures, while being able to synthesize multiple maps at the same time.</p>
<p>Our method can be improved in several ways. First, the adversarial expansion framework, while powerful, it has some potential pitfalls that hinder its widespread applicability. The same neural architecture is used for every texture but, as discussed, different choices on the architecture make different assumptions on the nature of the textures. We have proposed a generic architecture that works well for many examples, but recent advances in Neural Architecture Search [<xref ref-type="bibr" rid="CIT270">Mel+21</xref>] <target target-type="page" id="pges_92"/>may provide better priors on the optimal neural architecture to use for each sample. Furthermore, each texture synthesis network is trained from scratch. This is not only computationally costly, but learning to synthesize one texture may help in the synthesis of other textures, as shown by [<xref ref-type="bibr" rid="CIT248">Liu+20</xref>]. Fine-tuning pre-trained models for generating new textures may provide cheaper syntheses. Besides, our discriminator, while capable of detecting local artifacts, provides little control for separating such artifacts and global semantic errors. Recent work on image synthesis may provide guidance onto designing better discriminative models [<xref ref-type="bibr" rid="CIT354">SSK20</xref>], training procedures [<xref ref-type="bibr" rid="CIT441">Yun+19</xref>; <xref ref-type="bibr" rid="CIT365">Sin+21</xref>], image parametrizations [<xref ref-type="bibr" rid="CIT413">Wan+20a</xref>; <xref ref-type="bibr" rid="CIT172">HZ21</xref>; <xref ref-type="bibr" rid="CIT296">Nik+21</xref>] or perceptual loss functions [<xref ref-type="bibr" rid="CIT145">Hei+21</xref>].</p>
<p>Second, we proposed a synthesis solution based on manipulating latent spaces within the generative model, but explicitly training the network to generate tileable textures may provide better results than our approach. Besides, our sampling procedure could be extended for better selection of textures, by comparing histograms of the activations of the discriminator on the selected central area, instead of simply comparing minimum values.</p>
<p>Finally, our method has the advantage of being fully automatic, however, pre-processing the texture images so they are more easily tileable can help the synthesis process. For example, automatically rotating the textures so their repeating patterns are aligned with the axes was studied by [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>]. Powerful methods for artifact [<xref ref-type="bibr" rid="CIT062">Dek+15</xref>] and distortion [<xref ref-type="bibr" rid="CIT236">Li+19</xref>] removal could be applied as a pre-processing operation to the input textures before training the generative models, or as additional components to their loss function for improving tileability or homogeneity.</p>
</sec>
<sec id="c3.sA">
<label>3.A</label>
<title>Additional Details</title>
<p><bold>Training details:</bold> We use PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>] as our learning framework, and Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>] as our optimization algorithm. For it, we use an initial learning rate of &#x03B1; = 0.0002, which is divided by 5, after iterations 30000 and 40000. The networks are trained for a total of 50000 iterations, using a batch size of 1.The momentum parameters are kept to the default values of <inline-formula><mml:math id="M265" display='block'><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.9</mml:mn></mml:math></inline-formula> and <inline-formula><mml:math id="M266" display='block'><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.999</mml:mn></mml:math></inline-formula>, and <inline-formula><mml:math id="M267" display='block'><mml:mi>&#x03F5;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>8</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. We additionally apply an <inline-formula><mml:math id="M268" display='block'><mml:msub><mml:mo mathvariant="script">L</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> regularization of 10<sup>&#x2212;5</sup> to the optimization of both networks, which may help obtain better synthesized images [<xref ref-type="bibr" rid="CIT218">Kur+19</xref>]. This configuration is the same for the training of both the discriminator and the generator. In contrast to other work in unsupervised image generation [<xref ref-type="bibr" rid="CIT460">Zhu+17a</xref>], both networks are updated after each iteration. The models are trained and evaluated using a single NVIDIA GeForce GTX 1080Ti, leveraging half-precision training and automatic gradient scaling [<xref ref-type="bibr" rid="CIT274">Mic+18</xref>]. Evaluation is done in half precision for faster inference and reduced memory consumption. Using reduced numeric precision does not yield worse synthesized textures.</p>
<p><target target-type="page" id="pges_93"/><bold>Network architecture:</bold> We follow closely the architectures defined in [<xref ref-type="bibr" rid="CIT458">Zho+18</xref>]. For the <bold>Generator</bold> <inline-formula><mml:math id="M269" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula>, we use replication padding as the padding method for the entire architecture, ReLU non-linearities, and Instance Normalization [<xref ref-type="bibr" rid="CIT396">UVL16</xref>] before each non-linearity. This network receives textures of <inline-formula><mml:math id="M270" display='block'><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:math></inline-formula> pixels, which are first processed by an encoder <inline-formula><mml:math id="M271" display='block'><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula>, which transforms it to a <inline-formula><mml:math id="M272" display='block'><mml:mn>256</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>k</mml:mi><mml:mn>4</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>k</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:math></inline-formula> latent vector x<sub>0</sub>. Following standard practice on residual convolutional networks [<xref ref-type="bibr" rid="CIT139">He+16</xref>], the first layer of <inline-formula><mml:math id="M273" display='block'><mml:mi mathvariant="script">E</mml:mi></mml:math></inline-formula> is comprised by 64 7 &#x00D7; 7 convolutional filters. The rest of the network, unless stated otherwise, uses standard 3 &#x00D7; 3 kernels. This latent vector is further processed by 5 residual blocks <inline-formula><mml:math id="M274" display='block'><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">&#x2190;</mml:mo><mml:msub><mml:mi mathvariant="script">R</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, which maintain the latent vector dimensions. Then, a decoder <bold>G</bold> takes x<sub>5</sub> and, using three convolutional layers with transposed convolutions [<xref ref-type="bibr" rid="CIT254">LSD15</xref>], followed by ReLU operations, it returns the estimated texture, with <inline-formula><mml:math id="M275" display='block'><mml:mn>2</mml:mn><mml:mi mathvariant="normal">k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi mathvariant="normal">k</mml:mi></mml:math></inline-formula> resolution. As in the encoder, the last layer of <bold>G</bold> contains 64 7 &#x00D7; 7 convolutional filters. There are as many decoders <bold>G</bold> as individual maps in the texture stack. When tiling the latent space x<sub>0</sub>, this doubles it spatial resolution, thus becoming a <inline-formula><mml:math id="M276" display='block'><mml:mn>256</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi mathvariant="normal">k</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi mathvariant="normal">k</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:math></inline-formula> latent vector. As such, once processed by the residual blocks and the decoders, the network outputs a <inline-formula><mml:math id="M277" display='block'><mml:mn>4</mml:mn><mml:mi mathvariant="normal">k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>4</mml:mn><mml:mi mathvariant="normal">k</mml:mi></mml:math></inline-formula> texture stack.</p>
<p>For the <bold>discriminator</bold> <inline-formula><mml:math id="M278" display='block'><mml:mi>D</mml:mi></mml:math></inline-formula>, we simply follow a <italic>PatchGAN</italic> architecture [<xref ref-type="bibr" rid="CIT460">Zhu+17a</xref>; <xref ref-type="bibr" rid="CIT176">Iso+17</xref>; <xref ref-type="bibr" rid="CIT230">Led+17</xref>; <xref ref-type="bibr" rid="CIT458">Zho+18</xref>], which outputs local estimations of the probability of the input image being real or generated. We use 4 blocks of layers, each comprised of, in order, a convolutional layer with 4 &#x00D7; 4 kernels, an Instance Normalization [<xref ref-type="bibr" rid="CIT396">UVL16</xref>] layer, and a Leaky ReLU non-linearity [<xref ref-type="bibr" rid="CIT260">MHN13</xref>]. We use a stride of 2 in the first and last layers to decrease the dimensions of the latent representations. The first block of layers contains 64 trainable kernels, which are doubled after each block. The last layer thus contains 512 trainable filters. In the last layer, we remove the normalization operation and substitute the non-linearity by a Sigmoid operation, to transform the output into the (0, 1) range. The discriminator receives <inline-formula><mml:math id="M279" display='block'><mml:mn>2</mml:mn><mml:mi mathvariant="normal">k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi mathvariant="normal">k</mml:mi></mml:math></inline-formula> resolution stacks, and outputs a <inline-formula><mml:math id="M280" display='block'><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi mathvariant="normal">k</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi mathvariant="normal">k</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:math></inline-formula> resolution estimation map.</p>
</sec>
<sec id="c3.sB">
<label>3.B</label>
<title>Additional Comparisons, Results and Ablations</title>
<p>In this section, we outline the implementation details used for our comparison with other tileable texture synthesis methods, which included Histogram-Preserving Blending by Deliot and Heitz [<xref ref-type="bibr" rid="CIT063">DH19</xref>], Graph-Cuts synthesis by Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT241">Li+20</xref>], Repeated Pattern Detection by Rodriguez-Pardo <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>], Periodic Spatial GAN by Bergmann <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT028">BJV17</xref>] and Self-Organizing Textures by Niklasson <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT296">Nik+21</xref>]. We also provide further results in <xref ref-type="fig" rid="fig-3.15">Figure 3.15</xref> and an ablation study on the size of the model architectures in <xref ref-type="fig" rid="fig-3.14">Figure 3.14</xref>.</p>
<fig id="fig-3.14">
<label>Figure 3.14:</label>
<caption><title>An ablation study on the influence of the number of layers in each model in the quality of the generated outputs. We evaluate each network using the same input, marked using a blue box. We compare evaluate the influence of using discriminators with <italic>L</italic> &#x2208; {3, 4, 5} layers, marked as <italic>D<sub>L</sub></italic> and generators with <italic>L</italic> &#x2208; {4, 5, 6} residual blocks, marked as <italic>G<sub>L</sub></italic>. We show the configuration we use for every experiment in this chapter using a green box. On the bottom of each image, we show the training time for 50000 iterations. As shown, using a discriminator with 3 layers yields repetitive large-scale patterns, due to its reduced receptive field. Using <italic>D<sub>L</sub></italic> = 5 yields marginal improvements with respect to D<sub>L</sub> = 4, with the cost of increased runtimes. In terms of G<sub>L</sub>, we find that there are no significant visual differences between using 5 or 6 residual blocks, but we favor the former due to its reduced training time.</title></caption>
<alt-text>An ablation study on the influence of the number of layers in each model in the quality of the generated outputs. We evaluate each network using the same input, marked using a blue box. We compare evaluate the influence of using discriminators with L &#x2208; {3, 4, 5} layers, marked as DL and generators with L &#x2208; {4, 5, 6} residual blocks, marked as GL. We show the configuration we use for every experiment in this chapter using a green box. On the bottom of each image, we show the training time for 50000 iterations. As shown, using a discriminator with 3 layers yields repetitive large-scale patterns, due to its reduced receptive field. Using DL = 5 yields marginal improvements with respect to DL = 4, with the cost of increased runtimes. In terms of GL, we find that there are no significant visual differences between using 5 or 6 residual blocks, but we favor the former due to its reduced training time.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.14.jpg"/>
</fig>
<fig id="fig-3.15">
<label>Figure 3.15:</label>
<caption><title>More examples on the capabilities of SeamlessGAN on generating multiple tileable textures from a given input. On the left, we show the input sample <inline-formula><mml:math id="M283" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> and four crops used for generating tileable maps. On the middle column, the outputs of the generative model are shown. On their right, the same textures are tiled 2 &#x00D7; 2 times for visualization purposes.</title></caption>
<alt-text>More examples on the capabilities of SeamlessGAN on generating multiple tileable textures from a given input. On the left, we show the input sample T and four crops used for generating tileable maps. On the middle column, the outputs of the generative model are shown. On their right, the same textures are tiled 2 &#x00D7; 2 times for visualization purposes.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.15.jpg"/>
</fig>
<p><target target-type="page" id="pges_94"/><bold>Histogram-Preserving Blending</bold> We use the implementation provided by the authors. The only hyperparameter to be changed is the blending border size, which is set to the default value of 33% for all the textures.</p>
<p><bold>Graph-Cuts synthesis</bold> To allow for comparison between single-map texture synthesis methods, we run this method to output only the albedo map, using the default values of <italic>&#x03BB;</italic><sub>d</sub> = 1, a ratio of 0.64, a crop ratio of 0.6 and a margin size of 72 pixels. The images are not rescaled to any input or target size, instead, we use their original resolution.</p>
<p><bold>Repeated Pattern Detection</bold> This paper proposed the use of some preprocessing methods for improving the regularity of the input images, which were not used for our comparison. We use their default hyperparameter values, including <inline-formula><mml:math id="M281" display='block'><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.65</mml:mn><mml:mo>,</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.8</mml:mn></mml:math></inline-formula>, and the deep perceptual loss weighting proposed in [<xref ref-type="bibr" rid="CIT107">GEB15a</xref>].</p>
<p><bold>Periodic Spatial GAN</bold> This method can handle both single-image and multiple image datasets. We are interested in the former, and, as such, each GAN is trained using only one texture sample. We train their model using Theano [<xref ref-type="bibr" rid="CIT332">Al-+16</xref>], and use their default configuration of a learning rate of 0.0002, a <inline-formula><mml:math id="M282" display='block'><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> = 0.5 as the momentum term for the optimization algorithm, a batch size of 25, and train their network for 10 epochs, each comprised of 25000 training steps. The network parameters and sizes are exactly as defined in the paper.</p>
<p><bold>Self-Organizing Textures</bold> This method is computationally expensive and, as such, it requires that the input textures are rescaled to 128 &#x00D7; 128 pixels. We train their method using Tensorflow [<xref ref-type="bibr" rid="CIT001">Aba+16</xref>] for 5000 iterations. We use the default training configuration, with a starting learning rate of 0.002, which is then reduced by half at iterations 2000 and 4000. We do not modify the image parametrization design or loss function in any way.<target target-type="page" id="pges_95"/><target target-type="page" id="pges_96"/></p>
<fig id="fig-3.16">
<label><target target-type="page" id="pges_97"/>Figure 3.16:</label>
<caption><title>Additional comparison with previous methods. From left to right, top to bottom: Input texture, the results of Histogram Blending, by Deloit <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT063">DH19</xref>], the GraphCuts algorithm by Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT241">Li+20</xref>], Texture Stationarization by Moritz <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>], Repeated Pattern Detection by Rodriguez-Pardo <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT341">Rod+19</xref>], PSGAN by Bergmann <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT028">BJV17</xref>], Self-Organizing Textures by Niklasson <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT296">Nik+21</xref>]; and ours. Outputs are tiled a similar number of times (at least twice in each dimension) for better visualization. Our method generally captures better the overall structure of the texture, while providing seamless and semantically coherent borders, for enhancing tileability. From left to right, we sample c = 27, 36, 29, 1, 17, and 7 input crops before obtaining a tileable texture.</title></caption>
<alt-text>Additional comparison with previous methods. From left to right, top to bottom: Input texture, the results of Histogram Blending, by Deloit et al. [DH19], the GraphCuts algorithm by Li et al. [Li+20], Texture Stationarization by Moritz et al. [Mor+17], Repeated Pattern Detection by Rodriguez-Pardo et al. [Rod+19], PSGAN by Bergmann et al. [BJV17], Self-Organizing Textures by Niklasson et al. [Nik+21]; and ours. Outputs are tiled a similar number of times (at least twice in each dimension) for better visualization. Our method generally captures better the overall structure of the texture, while providing seamless and semantically coherent borders, for enhancing tileability. From left to right, we sample c = 27, 36, 29, 1, 17, and 7 input crops before obtaining a tileable texture.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-3.16.jpg"/>
</fig>
</sec>
</body>
<back>
<fn-group>
<fn id="c3-fn1"><label>1</label> <p>Additional publication details and results are included on <bold>the project website</bold>.</p></fn>
</fn-group>
</back>
</book-part>
<book-part id="c4" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_98"/><target target-type="page" id="pges_99"/>Chapter 4.</label>
<title>Neural Fields for BTF Encoding and Transfer</title>
</title-group>
</book-part-meta>
<body>
<fig id="fig-4.1">
<label>Figure 4.1:</label>
<caption><title>On the left, renders generated with NeuBTF materials. On their right, a slice of the input training BTF and the guidance image onto which we propagate the BTF measurements.</title></caption>
<alt-text>On the left, renders generated with NeuBTF materials. On their right, a slice of the input training BTF and the guidance image onto which we propagate the BTF measurements.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.1.jpg"/>
</fig>
<p>Neural material representations are becoming a popular way to represent materials for rendering. They are more expressive than analytic models and occupy less memory than tabulated BTFs. However, existing neural materials are immutable, meaning that their output for a certain query of UVs, camera, and light vector is fixed once they are trained. While this is practical when there is no need to edit the material, it can become very limiting when the fragment of the material used for training is too small or not tileable, which frequently happens when the material has been captured with a gonioreflectometer. In this chapter, we propose a novel neural material representation that jointly tackles the problems of BTF compression, tiling, and extrapolation. At test time, our method uses a guidance image as input to condition the neural BTF to the structural features of this input image. Then, the neural BTF can be queried as a regular BTF using UVs, camera, and light vectors. Every component in our framework is purposefully designed to maximize BTF encoding quality at minimal parameter count and computational complexity, achieving competitive compression rates compared with previous work. We demonstrate the results of our method on a variety of synthetic and captured materials, showing its generality and capacity to learn to represent many optical properties. The contributions in this chapter led to the following publication, currently under review:</p>
<fig id="fig-4">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.jpg"/>
</fig>
<disp-quote>
<p><target target-type="page" id="pges_100"/>&#x201C;NeuBTF: Neural Fields for BTF Encoding and Transfer&#x201D;</p>
<p>Carlos Rodriguez-Pardo, Konstantinos Kazatzis, Jorge Lopez-Moreno, Elena Garces</p>
<p><italic>Under Review</italic> (2023)</p>
</disp-quote>
<sec id="c4-s1">
<label>4.1</label>
<title>Introduction</title>
<p>A common approach to modeling real-world spatially-varying materials in computer graphics is through the use of Bidirectional Texture Functions (BTFs). This type of representation models the dense optical response of the material, and is more general than analytic representations such as a microfacet SVBRDF. However, BTFs typically occupy large amounts of memory. Recently, neural material representations are being proposed as an learning-based alternative to tabulated BTFs, providing a more compact solution while keeping the flexibility and generality of BTFs.</p>
<p>Creating digital representations of real material samples requires using an optical capture device, such as a gonioreflectometer [<xref ref-type="bibr" rid="CIT103">Gar+23</xref>], a smartphone [<xref ref-type="bibr" rid="CIT065">Des+18</xref>; <xref ref-type="bibr" rid="CIT066">Des+19</xref>; <xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT148">Hen+21</xref>], or a flatbed scanner [<xref ref-type="bibr" rid="CIT337">Rod+23a</xref>]. During the process, several choices must be made. First, it is important to select a patch of the material that contains enough spatial variability. Second, a process &#x015B;automatic or manual&#x015B; must be found to produce a tileable material that can be used to create seamless 3D renders. Finally, resources must be allocated for storage as needed. Making these choices when dealing with implicit or tabulated representations, such as in BTFs or neural materials, is particularly crucial. Once these representations are trained or captured, they cannot be easily modified and it is only possible to query them using the UVs, light, and camera vectors.</p>
<p>In this chapter, we propose a novel neural material representation that addresses these issues. Unlike existing neural approaches that are immutable once trained [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>; <xref ref-type="bibr" rid="CIT220">Kuz+21</xref>; <xref ref-type="bibr" rid="CIT316">Rai+20</xref>; <xref ref-type="bibr" rid="CIT221">Kuz+22</xref>], our model can be queried a test time with a guidance image that conditions the neural BTF to the structure provided by the guidance image. Our approach resembles synthesis by example and procedural processes, and can be used to extrapolate BTFs to large material samples, as well as to easily create tileable ones. Furthermore, our method achieve better compression rates than previous work on neural BTF representations.</p>
<p>To achieve this, we present a novel deep learning method that works at two steps. In the first step, we condition the neural BTF using a <italic>guidance image</italic> <target target-type="page" id="pges_101"/>as input. To this end, we use an <italic>autoencoder</italic> that outputs a high-dimensional latent representation of the material, a <italic>neural texture</italic>, which jointly encodes reflectance and structural properties. In the second step, the UV position of the latent representation, along with the camera and light vectors, are decoded by a fully-convolutional sinusoidal decoder [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>], a neural <italic>renderer</italic> to obtain the RGB values. Using a single BTF as input, we train the network end-to-end using a custom training procedure, loss function, and data augmentation policy. This policy, inspired by recent work on attribute transfer [<xref ref-type="bibr" rid="CIT338">RG21</xref>] (<xref ref-type="book-part" rid="c2">Chapter 2</xref>), allows the autoencoder to encode the relationship between structural features and reflectance, enabling the propagation of the BTF to novel input guidances. Once trained, the novel input guidances may come from the same material, a different material, or a structural pattern. An input guidance of the same material can be used to extrapolate the BTF to larger samples or create tileable BTFs, provided the input guidance is tileable. If the input guidance is a structural pattern, the local features can be used to synthesize novel materials.</p>
<p>In summary, we propose the following contributions:</p>
<list list-type="bullet">
<list-item><p>The first neural BTF representation with conditional input that can be used to extrapolate BTF measurements, easily create tileable BTFs, and synthesize novel materials.</p></list-item>
<list-item><p>We show how to leverage our system for rendering large-scale and tileable neural BTF generation using measurements captured with small portions of the material.</p></list-item>
<list-item><p>We demonstrate that our method works with synthetic and captured materials of diverse optical properties, including colored specular or anisotropy.</p></list-item>
</list>
</sec>
<sec id="c4-s2">
<label>4.2</label>
<title>Related Work</title>
<p>An accurate method for representing the optical properties of materials is through Bidirectional Texture Functions (BTFs) [<xref ref-type="bibr" rid="CIT057">Dan01</xref>]. BTFs are 6D functions that characterize all possible combinations of incoming and outgoing light and camera directions for the 2D spatial extent of a material. Although they are successful in representing materials, they have a major drawback in terms of memory requirements. Therefore, BTF compression has been a major research topic [<xref ref-type="bibr" rid="CIT091">FH08</xref>]. Non-neural approaches used dimensionality reduction techniques such as Principal Component Analysis (PCA) [<xref ref-type="bibr" rid="CIT212">Kou+03</xref>; <xref ref-type="bibr" rid="CIT282">MMK03</xref>; <xref ref-type="bibr" rid="CIT422">WGK14</xref>], vector quantization [<xref ref-type="bibr" rid="CIT137">HFM10</xref>], or clustering [<xref ref-type="bibr" rid="CIT390">Ton+02</xref>]. However, these approaches were recently surpassed by neural models [<xref ref-type="bibr" rid="CIT155">HS06</xref>] due to their flexibility and superior capacity to learn non-linear functions.</p>
<p><target target-type="page" id="pges_102"/><bold>Neural BTFs</bold> Rainer <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>] proposed the first method to use deep autoencoders to compress BTFs, surpassing PCA [<xref ref-type="bibr" rid="CIT422">WGK14</xref>] on captured BTFs. However, this approach required training a single neural network per material. To address this limitation, a later work by the same authors [<xref ref-type="bibr" rid="CIT316">Rai+20</xref>] proposed a generalization of this idea in which a single network was able to generalize to a variety of materials. Although these methods were very effective for compressing flat materials, they had some limitations when it came to modeling materials with volume. In their work, Kuznetsov <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>] improved the quality of neural materials by introducing a neural offset module that captures parallax effects. Further, their method also allowed for level-of-detail though MIP mapping by training a multi-resolution neural representation. However, grazing angles and silhouette effects remained a challenge for this approach. In a subsequent work, Kuznetsov <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT221">Kuz+22</xref>] explicitly trained the network using queries that span surface curvatures, effectively handling these cases. Representing fur, fabrics, and grass with neural reflectance fields was explored by Baatz <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT019">Baa+22</xref>] who proposed a representation that jointly models reflectance and geometry</p>
<p><target target-type="page" id="pges_103"/>All of these approaches share the idea of querying neural material using UVs, camera, and lighting vectors, but do not provide any functionality for modifying the material once the network is trained. In contrast, our approach can take a guidance image as input, which conditions the output to generate material variability.</p>
<p><bold>Material Synthesis and Tiling</bold> Texture synthesis is a long-standing problem in the field of computer graphics. The goal is to reconstruct a larger image given a small sample, leveraging the structural content and internal statistics of the input image. This concept has been used for synthesizing single images, BTFs, and full material models. For images, the most common strategies include PatchMatch [<xref ref-type="bibr" rid="CIT070">Dia+15</xref>], texture transport [<xref ref-type="bibr" rid="CIT005">AWL15</xref>], point processes [<xref ref-type="bibr" rid="CIT123">Gue+20</xref>; <xref ref-type="bibr" rid="CIT231">LH06</xref>], or neural networks [<xref ref-type="bibr" rid="CIT085">EM17</xref>; <xref ref-type="bibr" rid="CIT458">Zho+18</xref>; FAW19; RG22]. BTF synthesis, however, has received less attention. Steinhausen et al. [<xref ref-type="bibr" rid="CIT375">Ste+15a</xref>; <xref ref-type="bibr" rid="CIT376">Ste+15b</xref>] extrapolated</p>
<p> BTF captures to larger material samples using non-neural texture synthesis methods. For full materials, Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT247">LPG19</xref>] captured the appearance of materials by first estimating their BRDF and then synthesizing the high-resolution micro-structure from a dataset of measured SVBRDFs. Nagano <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT289">Nag+15</xref>] measured microscopic patches of the skin and used a convolutional filter to propagate the measurements to a spatially-varying texture. Deschaintre <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT067">DDB20</xref>] used an autoencoder to propagate SVBRDFs to large material samples. Also recently, Rodriguez-Pardo and Garces [<xref ref-type="bibr" rid="CIT338">RG21</xref>] propagated any kind of visual attribute having a single image as guidance. Their approach shares some similarity to ours, although they transfer 2D image attributes, while we transfer the full BTF.</p>
<p>Procedural models [<xref ref-type="bibr" rid="CIT164">Hu+22b</xref>; <xref ref-type="bibr" rid="CIT455">Zho+23</xref>; <xref ref-type="bibr" rid="CIT454">Zho+22</xref>] are nowadays very successful for generating tileable materials. Thanks to the use of a tileable template, these methods adjust the generated image to the features available in the template. As we show, our approach can also work with a binary template as input. However, guaranteeing predictable outputs given this kind of input is out of the scope of our technique, which can transfer BTF measurements having as input a guidance image of the same material.</p>
<p><bold>Other Neural Representations in Rendering</bold> Limited to BRDFs, neural networks trained with adaptive angular sampling have been explored to enable importance sampling [<xref ref-type="bibr" rid="CIT380">Szt+21</xref>], needed for Monte Carlo integration. Deep latent representations also allow for BRDF editions. For example, Hu <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT159">Hu+20</xref>] demonstrate that autoencoders can outperform classic PCA for the purpose of editing. Other applications of neural encodings in rendering are numerous. For instance, they have been used for scene prefiltering [<xref ref-type="bibr" rid="CIT020">BSK23</xref>], where geometry and materials are simplified to accommodate the LoD of the scene using a voxel-based representation and trained latent encodings. <target target-type="page" id="pges_104"/>For anisotropic microfacets, Gauthier <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT112">Gau+22</xref>] propose a cascaded architecture able to adjust the material parameters to the MIP mapping level. Encoding light transport using neural networks for real-time global illumination has also been explored [<xref ref-type="bibr" rid="CIT315">Rai+22</xref>; <xref ref-type="bibr" rid="CIT101">GMX22</xref>], showcasing promising results.</p>
</sec>
<sec id="c4-s3">
<label>4.3</label>
<title>Method</title>
<p>We present an overview of our approach in <xref ref-type="fig" rid="fig-4.2">Figure 4.2</xref>, where we show our inference and training pipelines. Our goal is twofold: First, find a compact representation for a BTF through the use of neural networks. Second, enable the extrapolation of the BTF according to guidance images used as input. In <xref ref-type="sec" rid="c4-s3.s1">Section 4.3.1</xref> we describe our inference pipeline and neural network, and in <xref ref-type="sec" rid="c4-s3.s2">Section 4.3.2</xref> our training process. <xref ref-type="sec" rid="c4-s4">Section 4.4</xref> contains specific implementation and design details of the neural networks.</p>
<fig id="fig-4.2">
<label>Figure 4.2:</label>
<caption><title>An overview of our neural BTF inference and training processes. Top-Inference: Using a guidance image <inline-formula><mml:math id="M284" display='block'><mml:mi mathvariant="script">G</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi mathvariant="normal">H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi mathvariant="normal">W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, we use our trained autoencoder to generate the Neural Texture <inline-formula><mml:math id="M285" display='block'><mml:mi mathvariant="script">A</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="script">T</mml:mi><mml:mi mathvariant="script">G</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi mathvariant="normal">H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi mathvariant="normal">W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi mathvariant="normal">D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, which preserves the spatial resolution of the input image but represents a higher-dimensional learned representation. This Neural Texture <inline-formula><mml:math id="M286" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">T</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, along with the trained renderer <inline-formula><mml:math id="M287" display='block'><mml:mi mathvariant="script">R</mml:mi></mml:math></inline-formula>, can be queried as a regular BTF, using UVs, and target camera and light positions for regular rendering. Bottom-Train: During training, we randomly select an input <italic>view</italic> <inline-formula><mml:math id="M288" display='block'><mml:msub><mml:mi mathvariant="normal">V</mml:mi><mml:mrow><mml:msub><mml:mover><mml:mi>&#x03C9;</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mover><mml:mi>&#x03C9;</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> and a target <inline-formula><mml:math id="M289" display='block'><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>. To both views, we apply random rescale and cropping. Then, only to <inline-formula><mml:math id="M290" display='block'><mml:msub><mml:mi mathvariant="normal">V</mml:mi><mml:mrow><mml:msub><mml:mover><mml:mi>&#x03C9;</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mover><mml:mi>&#x03C9;</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>, we randomly apply hue variations, gaussian blur, and noise, and feed it to the <italic>autoencoder</italic>, which returns a 2D <italic>latent representation</italic> of the material. A fully-convolutional <italic>decoder</italic> with sinusoidal activations receives both this latent space and the target <inline-formula><mml:math id="M291" display='block'><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub></mml:math></inline-formula> camera and light angles, and estimates <inline-formula><mml:math id="M292" display='block'><mml:msub><mml:mover><mml:mi mathvariant="normal">V</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>. This output is compared with <inline-formula><mml:math id="M293" display='block'><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> using a multifaceted loss function.</title></caption>
<alt-text>An overview of our neural BTF inference and training processes. Top-Inference: Using a guidance image G&#x2208;&#x211D;H&#x00D7;W&#x00D7;3, we use our trained autoencoder to generate the Neural Texture A(G)=TG&#x2208;&#x211D;H&#x00D7;W&#x00D7;D, which preserves the spatial resolution of the input image but represents a higher-dimensional learned representation. This Neural Texture TG, along with the trained renderer R, can be queried as a regular BTF, using UVs, and target camera and light positions for regular rendering. Bottom-Train: During training, we randomly select an input view V&#x03C9;&#x007E;o,&#x03C9;&#x007E;i and a target V&#x03C9;o,&#x03C9;i. To both views, we apply random rescale and cropping. Then, only to V&#x03C9;&#x007E;o,&#x03C9;&#x007E;i, we randomly apply hue variations, gaussian blur, and noise, and feed it to the autoencoder, which returns a 2D latent representation of the material. A fully-convolutional decoder with sinusoidal activations receives both this latent space and the target &#x03C9;o,&#x03C9;i camera and light angles, and estimates V^&#x03C9;o,&#x03C9;i. This output is compared with V&#x03C9;o,&#x03C9;i using a multifaceted loss function.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.2.jpg"/>
</fig>
<sec id="c4-s3.s1">
<label>4.3.1</label>
<title>Inference</title>
<p>Our neural network is composed of three modules: an <italic>autoencoder</italic> <inline-formula><mml:math id="M294" display='block'><mml:mi mathvariant="script">A</mml:mi></mml:math></inline-formula>, a <italic>neural texture</italic> <inline-formula><mml:math id="M295" display='block'><mml:mi mathvariant="script">T</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, and a <italic>renderer</italic> <inline-formula><mml:math id="M296" display='block'><mml:mi mathvariant="script">R</mml:mi></mml:math></inline-formula>. The renderer <inline-formula><mml:math id="M297" display='block'><mml:mi mathvariant="script">R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="script">T</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">u</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">v</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>RGB</mml:mi></mml:math></inline-formula> takes as input the feature vector at the (u, v) coordinates of the neural texture <inline-formula><mml:math id="M298" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> , the view <inline-formula><mml:math id="M299" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>O</mml:mi></mml:msub></mml:math></inline-formula> and light <inline-formula><mml:math id="M300" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> positions, and returns an RGB value. <inline-formula><mml:math id="M301" display='block'><mml:mi mathvariant="script">R</mml:mi></mml:math></inline-formula> acts as a conventional BTF and can be used as such in any render engine. An input <italic>guidance</italic> image, <inline-formula><mml:math id="M302" display='block'><mml:mi mathvariant="script">G</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, is used during inference to condition the generation of the <italic>neural texture</italic> <inline-formula><mml:math id="M303" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula>. This conditioning allows us to propagate the <target target-type="page" id="pges_105"/>learned reflectance to novel guidance images that can be: a larger sample of the same material, a different material, or a structural image. In the simpler case, the guidance image comes from the BTF used for training, and our process is equivalent to previous work [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>; <xref ref-type="bibr" rid="CIT220">Kuz+21</xref>].</p>
<p>The autoencoder <inline-formula><mml:math id="M304" display='block'><mml:mi mathvariant="script">A</mml:mi></mml:math></inline-formula> takes the guidance image <inline-formula><mml:math id="M305" display='block'><mml:mi mathvariant="script">G</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> as input and outputs a neural texture <inline-formula><mml:math id="M306" display='block'><mml:mi mathvariant="script">T</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> with the same size <inline-formula><mml:math id="M307" display='block'><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula> as the input guidance image but with more latent dimensions <italic>D</italic>. As a result, each pixel in the guidance image has a higher-dimensional neural representation in <inline-formula><mml:math id="M308" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula>. Because it is trained without explicit supervision, this latent representation can capture the reflectance and structural patterns automatically. Kuznetsov <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>] also used a latent neural representation of the material, however, lacking the initial autoencoder their approach cannot synthesize novel BTFs without retraining, while our conditioning module allow us to generate novel BTFs during test time. The autoencoder and the renderer are neural networks trained jointly, using an end-to-end image-to-image approach we describe below.</p>
</sec>
<sec id="c4-s3.s2">
<label>4.3.2</label>
<title>Training</title>
<p><xref ref-type="fig" rid="fig-4.2">Figure 4.2</xref> (bottom) illustrates our training process. It has two objectives: First, equivalent to regular BTF encoding, we aim to find the mapping between camera direction <inline-formula><mml:math id="M309" display='block'><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, light direction <italic>w</italic><sub>i</sub>, and output slices of the BTF: <inline-formula><mml:math id="M310" display='block'><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:msub><mml:mi mathvariant="normal">V</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>BTF</mml:mi></mml:math></inline-formula>. Second, we aim to condition the synthesis process with an input guidance image, <inline-formula><mml:math id="M311" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula>.</p>
<p>To this end, for training, we feed the model with images, <inline-formula><mml:math id="M312" display='block'><mml:msub><mml:mi mathvariant="normal">V</mml:mi><mml:mrow><mml:msub><mml:mover><mml:mi>&#x03C9;</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mover><mml:mi>&#x03C9;</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>, which are randomly sampled from the BTF, and are subject to additional data augmentation processes.</p>
<p>This extensive augmentation process guarantees invariance to different input variations during test time, like camera or illumination conditions, while keeps consistency of the outputs.</p>
<p><bold>Loss Function Design</bold> Our loss function, that compares ground truth slices <inline-formula><mml:math id="M313" display='block'><mml:msub><mml:mi mathvariant="normal">V</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> with generated ones <inline-formula><mml:math id="M314" display='block'><mml:mi mathvariant="script">R</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="script">A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">V</mml:mi><mml:mrow><mml:msub><mml:mover><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mover><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover><mml:mi mathvariant="normal">i</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03C9;</mml:mi><mml:mi mathvariant="normal">i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, is a weighted sum of three terms: a pixel-wise loss, a style loss, and a frequency loss,</p>
<p><disp-formula id="Eq004-1"><label>(4.1)</label> <mml:math id="M315" display='block'><mml:mi mathvariant="script">L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>freq</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>freq</mml:mtext></mml:msub></mml:math></disp-formula></p>
<p>The main driver of our loss is the pixel-wise norm <inline-formula><mml:math id="M316" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> produces sharper results than higher-order alternatives, such as <inline-formula><mml:math id="M317" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> [<xref ref-type="bibr" rid="CIT176">Iso+17</xref>; <xref ref-type="bibr" rid="CIT338">RG21</xref>]. Following [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>], we apply a &#x03B8;o(x + 1) compression to improve the model results on high dynamic range. This compression is only done to the pixel-wise component of the loss function. Inspired by recent work on texture synthesis, capture <target target-type="page" id="pges_106"/>and transfer [<xref ref-type="bibr" rid="CIT263">Mar+20</xref>; <xref ref-type="bibr" rid="CIT004">AWL13</xref>; <xref ref-type="bibr" rid="CIT339">RG22</xref>; <xref ref-type="bibr" rid="CIT337">Rod+23a</xref>; <xref ref-type="bibr" rid="CIT454">Zho+22</xref>], we introduce a <inline-formula><mml:math id="M318" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>style</mml:mtext></mml:msub></mml:math></inline-formula> loss to help the model generate higher quality and sharper results. Further, to mitigate the spectral bias of convolutional neural networks and help ameliorate the results further, we also introduce a <italic>focal frequency loss</italic> into our learning framework [<xref ref-type="bibr" rid="CIT183">Jia+21</xref>]. This combination of loss functions proves effective for our problem, without the need for complex adversarial losses which could reduce efficiency or destabilize training.</p>
<p><bold>Data Augmentation</bold> We train our models using a comprehensive data augmentation policy aimed at achieving high quality reflectance propagation, increasing performance and generalization, and allowing for generation of multiple resolution materials at test time. We build upon recent work on material transfer [<xref ref-type="bibr" rid="CIT338">RG21</xref>] and use images of the material taken under different illumination and viewing conditions as inputs to our autoencoder. This helps it generalize to novel capture setups, which allows for multiple applications we describe on <xref ref-type="sec" rid="c4-s6">Section 4.6</xref>. In particular, we use every image available on the input BTF, selected uniformly at random for each element in each batch during training. As in [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>; <xref ref-type="bibr" rid="CIT338">RG21</xref>; Rod+23a],</p>
<p>we also use random rescaling, which helps the model generalize to new scales, and, and build neural materials of multiple resolutions at test time, as we describe on <xref ref-type="sec" rid="c4-s6.s3">Section 4.6.3</xref> inspired by recent work on image synthesis [<xref ref-type="bibr" rid="CIT338">RG21</xref>; Tex+20b; RG22], we use random cropping, which helps generalization by effectively increasing the dataset size. Finally, we extend the color augmentation policy in [<xref ref-type="bibr" rid="CIT338">RG21</xref>] with <target target-type="page" id="pges_107"/>random hue changes across the entire color wheel, and introduce random Gaussian noise and blurs to the input images, to help it generalize further, as proposed in [<xref ref-type="bibr" rid="CIT337">Rod+23a</xref>] (<xref ref-type="book-part" rid="c5">Chapter 5</xref>).</p>
</sec>
</sec>
<sec id="c4-s4">
<label>4.4</label>
<title>Implementation details</title>
<p>Our model size, function weighting, optimizer and training configuration, and data augmentation hyperparameters were selected using a combination of manual tuning and Bayesian hyperparameter optimization using <italic>Weights and Biases</italic> [<xref ref-type="bibr" rid="CIT032">Bie20</xref>].</p>
<p><bold>Autoencoder</bold> For the <italic>autoencoder</italic>, we use a lightweight U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] with a few modifications to tailor it for our problem. Inspired by recent work on CNN design, we leverage ConvNext [<xref ref-type="bibr" rid="CIT253">Liu+22b</xref>] Blocks across our model, with depth-wise convolutions using 5 &#x00D7; 5 kernels. We empirically observe that ConvNext blocks achieve higher quality structural editions at a lower parameter count than vanilla U-Net blocks. To further help convergence and preserve details in the input images, we use residual connections [<xref ref-type="bibr" rid="CIT139">He+16</xref>; <xref ref-type="bibr" rid="CIT069">Dia+20</xref>; <xref ref-type="bibr" rid="CIT337">Rod+23a</xref>] in every convolutional block of the model. We use 1 &#x00D7; 1 convolutions on the skip connections and residual scaling [<xref ref-type="bibr" rid="CIT194">Kar+20a</xref>]. As in [<xref ref-type="bibr" rid="CIT253">Liu+22b</xref>], we use Layer Normalization [<xref ref-type="bibr" rid="CIT018">BKH16</xref>] and GELU non-linearities [<xref ref-type="bibr" rid="CIT147">HG16</xref>]. On the bottleneck of the model, we introduce an attention module [<xref ref-type="bibr" rid="CIT425">Woo+18</xref>] to help the model learn longer-range dependencies. To avoid checkerboard artifacts [<xref ref-type="bibr" rid="CIT298">ODO16</xref>], we use nearest neighbors interpolation for upsampling. Inspired by recent work on tileable material generation [<xref ref-type="bibr" rid="CIT454">Zho+22</xref>], we use <italic>circular padding</italic> throughout the model. We initialize its weights using <italic>orthogonal initialization</italic> [<xref ref-type="bibr" rid="CIT013">ACB19</xref>], which helps avoid exploding gradients.</p>
<p><bold>Renderer</bold> For the <italic>renderer</italic>, we build upon SIREN [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>] MLPs, with additional modifications to enhance its performance for our problem. We use 1 &#x00D7; 1 convolutions instead of vanilla linear layers, to allow for end-to-end training using 2D images. Further, we introduce Layer Normalization [<xref ref-type="bibr" rid="CIT018">BKH16</xref>] before each sinusoidal non-linearity, which stabilizes training. Finally, inspired by [<xref ref-type="bibr" rid="CIT267">Meh+21</xref>], we use residual connections [<xref ref-type="bibr" rid="CIT139">He+16</xref>], to help preserve the information of the input vector across the decoder layers. Model weight initialization follows [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>]. With sinusoidal activations, we observe significantly higher reconstruction quality and training dynamics than with ReLU [<xref ref-type="bibr" rid="CIT290">NH10</xref>] MLPs, which are common for BTF compression [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>; <xref ref-type="bibr" rid="CIT316">Rai+20</xref>; <xref ref-type="bibr" rid="CIT220">Kuz+21</xref>]. Because the network is fully-convolutional, it can take as input feature vectors of any size. This is very convenient for our use cases when the input guidance image have a size different from the size of the original BTF used to train it.</p>
<p><bold>Data Generation</bold> We generate synthetic ground truth BTF data using a proprietary CPU path-tracer based on the architectures of PBRT [<xref ref-type="bibr" rid="CIT307">PJH20</xref>], <target target-type="page" id="pges_108"/>Mitsuba [<xref ref-type="bibr" rid="CIT297">Nim+19</xref>] and Wave-front [<xref ref-type="bibr" rid="CIT226">LKA13</xref>]. The scene is set up with a single plane-mesh, and a directional light with unit irradiance rotating around the surface. We use a custom camera for skewed configurations to avoid post processing. We render a total of 81 &#x00D7; 81 = 6561 HDR images, changing the camera and lighting angles for each, following [<xref ref-type="bibr" rid="CIT353">SSK03</xref>]. The data are generated with 16 samples per BTF texel using a 10-core CPU, which takes approximately one hour for 512 &#x00D7; 512 resolution materials. For the rendered SVBRDFs, we use the material model in [<xref ref-type="bibr" rid="CIT103">Gar+23</xref>] (<xref ref-type="sec" rid="c2-sC">Chapter 2.C</xref>), without the transmittance and opacity maps. We use both procedurally and scanned SVBRDFs to generate our sythetic BTFs.</p>
<p>For acquired BTFs, we test on a variety of publicly available datasets, including <italic>UBO 2003</italic> [<xref ref-type="bibr" rid="CIT353">SSK03</xref>], <italic>UBO 2014</italic> [<xref ref-type="bibr" rid="CIT327">Rei+14</xref>], <italic>Atrium</italic>, <italic>UTIA 2012</italic> [<xref ref-type="bibr" rid="CIT131">Hai12</xref>], or <italic>UTIA MAM 2014</italic> [<xref ref-type="bibr" rid="CIT090">Fil+18</xref>]. Each of those has their own spatial and angular resolutions, we refer the reader to their individual specifications and to the Survey in [<xref ref-type="bibr" rid="CIT158">HS13</xref>] for additional details. We use <italic>btf-extractor</italic> to read the acquired BTF files. Our method works seamlessly with both measured and synthetically generated data without the need to change neither the neural module or the training procedure.</p>
<p><bold>Optimization</bold> We use PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>] and Torchvision [<xref ref-type="bibr" rid="CIT262">MR10</xref>] for training. We train our models for 300 epochs, using a batch size of 80, and an input crop size of 128 &#x00D7; 128 pixels, as in [<xref ref-type="bibr" rid="CIT338">RG21</xref>; <xref ref-type="bibr" rid="CIT339">RG22</xref>; <xref ref-type="bibr" rid="CIT337">Rod+23a</xref>]. We set an initial learning rate of 0.0015, which we divide by 2 every 100 iterations using a step scheduler. We empirically observe training instabilities using Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>] on early iterations. <target target-type="page" id="pges_109"/>To mitigate this, we use RAdam [<xref ref-type="bibr" rid="CIT249">Liu+19a</xref>], which significantly stabilizes training. We also perform gradient norm clipping [<xref ref-type="bibr" rid="CIT445">Zha+20</xref>] (set to <inline-formula><mml:math id="M319" display='block'><mml:mtext>maxnorm</mml:mtext><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula>) and an <inline-formula><mml:math id="M320" display='block'><mml:mi>&#x03F5;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. We use a Dropout [<xref ref-type="bibr" rid="CIT373">Sri+14</xref>] rate of 0.1 on the autoencoder and a weight decay of 10<sup>&#x2212;6</sup>, to avoid overfitting. To further regularize the models, we leverage mixed precision training and automatic scaling [<xref ref-type="bibr" rid="CIT274">Mic+18</xref>]. This training process takes approximately 3 hours per material on a Nvidia RTX 3060 GPU. These training times are constant regardless of the material spatial resolution. For comparison, our model takes around four times longer to train than previous work [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>], because we need to optimize the autoencoder parameters and our loss function is multifaceted. Convergence is typically achieved on 100-150 iterations, longer training only helps resolve additional details. A subset of the training set is also leveraged as a validation set, which helps track the model performance on larger images. We use this validation dataset for early stopping regularization. During test time, we evaluate the model using <italic>half precision</italic>.</p>
<p><bold>Loss Function</bold> We weight each part of the loss function as: <inline-formula><mml:math id="M321" display='block'><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>0.25</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>freq</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, and <inline-formula><mml:math id="M322" display='block'><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:msub><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>. For <inline-formula><mml:math id="M323" display='block'><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>style</mml:mtext></mml:msub></mml:math></inline-formula>, we use the <italic>AlexNet</italic> variant of <italic>LPIPS</italic> [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>], which allows for a lightweight yet powerful texture regularizer [<xref ref-type="bibr" rid="CIT339">RG22</xref>; <xref ref-type="bibr" rid="CIT337">Rod+23a</xref>]. This metric expects input on the [&#x2212;1, 1] range, so we perform a simple 2x &#x2212; 1 transformation to the model target and output, which lie on the [0, 1] range before computing the loss.</p>
<p><bold>Data Augmentation</bold> In order, we apply to both input and target images: random rescales in the range of [0.25, 1.5] and random cropping. Then, only to the input image: random hue changes, random grayscale (<italic>p</italic> = 0.2), random blurs (<italic>p</italic> = 0.2, kernel size = 5) and random Gaussian noise (<italic>p</italic> = 0.2). We empirically observe that higher probabilities, or more complex policies like random erasing [<xref ref-type="bibr" rid="CIT452">Zho+20</xref>] were detrimental to the model performance for some materials. We use Kornia [<xref ref-type="bibr" rid="CIT334">Rib+20</xref>] for these augmentations.</p>
<p><bold>Model Sizes</bold> Our <italic>autoencoder</italic> has 4 hidden layers in the encoder, with a <italic>width factor</italic> of 5. All other model details follow [<xref ref-type="bibr" rid="CIT344">RFB15</xref>]. The outputs of the autoencoder (e.g. the 2D latent space) are regularized with Layer Normalization [<xref ref-type="bibr" rid="CIT018">BKH16</xref>]. We use <italic>D</italic> = 14 latent dimensions for our neural materials, following [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>]. For the <italic>decoder</italic>, we use 3 hidden layers with 40 neurons each, achieving similar model sizes as previous work on neural BRDF representations [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>; <xref ref-type="bibr" rid="CIT380">Szt+21</xref>]. Following [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>], introduce a ReLU non linearity [<xref ref-type="bibr" rid="CIT290">NH10</xref>] to the final layer of the decoder to avoid negative reflectance values.</p>
</sec>
<sec id="c4-s5">
<label>4.5</label>
<title>Evaluation</title>
<sec id="c4-s5.s1">
<label>4.5.1</label>
<title>Qualitative Analysis</title>
<p>We evaluate NeuBTF on materials from different sources including acquired BTFs from [<xref ref-type="bibr" rid="CIT422">WGK14</xref>; <xref ref-type="bibr" rid="CIT131">Hai12</xref>], and rendered BTFs from procedurally generated and <target target-type="page" id="pges_110"/>scanned SVBRDFs. In <xref ref-type="fig" rid="fig-4.3">Figure 4.3</xref>, we show examples of the results of NeuBTF for a variety of materials with highly complex structures and reflectance properties, like colored specular (first column) or anisotropy (last). We show some additional results in <xref ref-type="fig" rid="fig-4.4">Figure 4.4</xref> for materials of different datasets. As shown, our model achieves high quality reconstructions regardless on the type of data source.</p>
<fig id="fig-4.3">
<label>Figure 4.3:</label>
<caption><title>Some tileable neural materials achieved with our method. On the top row, we show a slice of the BTF used to train a NeuBTF representation. With the tileable guidance images shown on the middle row, we propagate the neural texture using our autoencoders. These neural textures can be rendered to generate realistic images (bottom). We highlight the training dataset surface area as a green inset.</title></caption>
<alt-text>Some tileable neural materials achieved with our method. On the top row, we show a slice of the BTF used to train a NeuBTF representation. With the tileable guidance images shown on the middle row, we propagate the neural texture using our autoencoders. These neural textures can be rendered to generate realistic images (bottom). We highlight the training dataset surface area as a green inset.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.3.jpg"/>
</fig>
<fig id="fig-4.4">
<label>Figure 4.4:</label>
<caption><title>Qualitative results on a variety of BTFs, from different sources. From left to right, we show results on the UBO [<xref ref-type="bibr" rid="CIT422">WGK14</xref>] BTF dataset, the UTIA [<xref ref-type="bibr" rid="CIT131">Hai12</xref>] BTF dataset and two synthetic materials rendered from Substance SVBRDFs.</title></caption>
<alt-text>Qualitative results on a variety of BTFs, from different sources. From left to right, we show results on the UBO [WGK14] BTF dataset, the UTIA [Hai12] BTF dataset and two synthetic materials rendered from Substance SVBRDFs.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.4.jpg"/>
</fig>
<p>In <xref ref-type="fig" rid="fig-4.5">Figure 4.5</xref>, we show a colored visualization for a few channels of the latent neural texture <inline-formula><mml:math id="M324" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> found for a variety of materials. Because the values for the neural texture are unbounded, to each channel <inline-formula><mml:math id="M325" display='block'><mml:mi>c</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula>, we standardize them to 0 mean and unit variance, and apply a <inline-formula><mml:math id="M326" display='block'><mml:mi>sigmoid</mml:mi><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula> non-linearity to make the maps comparable. Without any explicit training, the models learn to separate distinct parts of the material. For example, the model finds distinct latent spaces for <italic>warp</italic> and <italic>weft</italic> yarns on woven fabrics, or separation between color and geometric patterns. This disentanglement provides clues on why the material propagation is possible, and suggests potential future research directions for fine-grained neural material edition.</p>
<fig id="fig-4.5">
<label>Figure 4.5:</label>
<caption><title>A selection of latent channels learned by NeuBTF for a variety of materials. We use a colorspace to help visualization. Without explicit supervision, the model internally learns semantically meaningful latent spaces. For instance, in the second example on the left, the two leftmost latent spaces encode geometry, while the other two encode the two distinct colors of the printed pattern over the yarns.</title></caption>
<alt-text>A selection of latent channels learned by NeuBTF for a variety of materials. We use a colorspace to help visualization. Without explicit supervision, the model internally learns semantically meaningful latent spaces. For instance, in the second example on the left, the two leftmost latent spaces encode geometry, while the other two encode the two distinct colors of the printed pattern over the yarns.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.5.jpg"/>
</fig>
<p>Finally, we also evaluate whether the encoded materials preserve the geometric properties in their training datasets. We leverage <italic>Photometric Stereo</italic> [<xref ref-type="bibr" rid="CIT174">Ike81</xref>] to compute surface normals of encoded and ground truth materials. We provide results in <xref ref-type="fig" rid="fig-4.6">Figure 4.6</xref>, where we show that the reconstructed surface geometry is accurate albeit noisier, and preserves the structure of the material.</p>
<fig id="fig-4.6">
<label>Figure 4.6:</label>
<caption><title>Surface normals estimated using <italic>Photometric Stereo</italic> [<xref ref-type="bibr" rid="CIT174">Ike81</xref>] for ground truth materials and the materials reconstructed by NeuMat. As shown, the encoded materials preserve geometric information without explicit training.</title></caption>
<alt-text>Surface normals estimated using Photometric Stereo [Ike81] for ground truth materials and the materials reconstructed by NeuMat. As shown, the encoded materials preserve geometric information without explicit training.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.6.jpg"/>
</fig>
</sec>
<sec id="c4-s5.s2">
<label><target target-type="page" id="pges_111"/>4.5.2</label>
<title>Compression comparisons with previous work</title>
<p>In <xref ref-type="table" rid="c4-tab1">Table 4.1</xref>, we show the number of trainable parameters on the decoders of different neural BTF compression algorithms. As shown, our model is competitive with previous work in terms of trainable parameters. This is achieved as we use more complex loss functions than previous work, which help regularize the models, and because our sinusoidal MLP achieves higher quality reconstructions for natural signals than ReLU MLPs, as shown in [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>]. NeuMIP [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>] uses smaller MLPs, however, they require an additional decoder for their <italic>neural offset</italic> module, which helps them encode parallax effects (See <xref ref-type="fig" rid="fig-4.7">Figure 4.7</xref>), for which our model struggles. Our decoder has one order of magnitude fewer parameters than [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>; <xref ref-type="bibr" rid="CIT316">Rai+20</xref>], however, the method in [<xref ref-type="bibr" rid="CIT316">Rai+20</xref>] provides the benefit of fast encoding of new materials, while ours requires a different model for each new material.</p>
<table-wrap id="c4-tab1">
<label>Table 4.1:</label>
<caption><title>Number of trainable parameters in the decoders and amount of latent texture channels for different neural BTF compression methods, measured using Torchinfo [<xref ref-type="bibr" rid="CIT438">Yep20</xref>]. Exact comparisons are challenging, as [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>] optimizes a multi-level texture pyramid and [<xref ref-type="bibr" rid="CIT316">Rai+20</xref>] learns a latent vector which can encode novel materials. For neither our method nor [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>; <xref ref-type="bibr" rid="CIT316">Rai+20</xref>], we count the parameters in the encoders, as they are not needed for using the materials on rendering systems.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="right"><p>Method</p></th>
<th valign="top" align="center"><p><bold>NeuBTF (ours)</bold></p></th>
<th valign="top" align="center"><p>[<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>]</p></th>
<th valign="top" align="center"><p>[<xref ref-type="bibr" rid="CIT317">Rai+19</xref>]</p></th>
<th valign="top" align="center"><p>[<xref ref-type="bibr" rid="CIT316">Rai+20</xref>]</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="right"><p>Decoder Parameters</p></td>
<td valign="top" align="center"><p><bold>3011</bold></p></td>
<td valign="top" align="center"><p>3332</p></td>
<td valign="top" align="center"><p>35725</p></td>
<td valign="top" align="center"><p>38269</p></td>
</tr>
<tr>
<td valign="top" align="right"><p>Texture Channels</p></td>
<td valign="top" align="center"><p><bold>14</bold></p></td>
<td valign="top" align="center"><p><bold>14</bold></p></td>
<td valign="top" align="center"><p><bold>14</bold></p></td>
<td valign="top" align="center"><p>38</p></td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-4.7">
<label><target target-type="page" id="pges_112"/>Figure 4.7:</label>
<caption><title>A failure case of our method. Compared to NeuMIP [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>], which explictly models parallax effects, our model struggles to accurately encode materials with strong displacements, as this synthetic cable knit from Substance3D. For this type of materials, NeuBTF accurately encodes orthogonal viewing angles (top row), however, it struggles at grazing angles (bottom row).</title></caption>
<alt-text>A failure case of our method. Compared to NeuMIP [Kuz+21], which explictly models parallax effects, our model struggles to accurately encode materials with strong displace-ments, as this synthetic cable knit from Substance3D. For this type of materials, NeuBTF accurately encodes orthogonal viewing angles (top row), however, it struggles at grazing angles (bottom row).</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.7.jpg"/>
</fig>
</sec>
<sec id="c4-s5.s3">
<label>4.5.3</label>
<title>Limitations</title>
<p>As we show in <xref ref-type="fig" rid="fig-4.7">Figure 4.7</xref>, our model struggles with materials with strong displacement. While our method provides accurate encodings on viewing angles close to the material surface, it cannot accurately encode grazing angles for such extreme cases. Displacement maps translate the geometric position of the points over the surface, breaking the underlying assumptions behind our neural texture. NeuMIP [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>] solves this issue by explicitly modelling parallax effects with a <italic>neural offset</italic> module. While we did not observe that such extension was needed for acquired BTF data, like the UBO2014 [<xref ref-type="bibr" rid="CIT327">Rei+14</xref>] dataset, introducing a similar module into our editable neural material framework is an interesting future research direction to increase its generality.</p>
</sec>
</sec>
<sec id="c4-s6">
<label>4.6</label>
<title>Applications</title>
<sec id="c4-s6.s1">
<label>4.6.1</label>
<title>Reflectance Propagation and Tileable BTFs</title>
<p>Many material reflectance acquisition devices are limited in the surface dimensions they can digitize. This hinders their applicability to many real-world materials, which exhibit variations that cannot be captured at such small scales. Further, in many applications like SVBRDF acquisition, obtaining larger samples of the material improves realism and helps tileable texture synthesis. In this context, previous work on BTF reflectance compression inherits the surface area limitations of the capture devices used to generate their <target target-type="page" id="pges_113"/>training data. Our method can easily be applied for reflectance propagation. We build upon the work of Rodriguez-Pardo and Garces [<xref ref-type="bibr" rid="CIT338">RG21</xref>] and leverage our <italic>encoder</italic> to propagate the neural texture optimized using a small portion of the material (e.g. a 1 &#x00D7; 1 cm capture) to a larger portion of the same material, represented with a <italic>guidance image</italic> captured using a commodity device like a flatbed scanner. Because our model is trained using a large amount of lighting conditions, as in [<xref ref-type="bibr" rid="CIT338">RG21</xref>], the propagation is invariant to how the images are illuminated. We show results of such pipeline in <xref ref-type="fig" rid="fig-4.3">Figure 4.3</xref>. For instance, on the last row, we show an anisotropic and specular Silver Jacquard fabric, for which we generated a BTF by rendering a 1 &#x00D7; 1 cm SVBRDF. This small crop cannot represent the complex pattern in the fabric, which we show on the <italic>guidance image</italic>, which covers a 10 &#x00D7; 10 cm area. Using our encoder, we propagate the neural material to this guidance image, generating a new, high-resolution, latent space which we can render, enabling realistic material representations with a reduced digitization cost. This propagated neural material has a 2000 &#x00D7; 2000 texels resolution, and it requires no retraining during test time.</p>
<p>Relatedly, our propagation framework can also be easily leveraged for generating tileable BTFs. Given any <italic>guidance image</italic> of the material, we can generate a tileable version of it, either using manual editions by artists or automatic algorithms [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>; <xref ref-type="bibr" rid="CIT241">Li+20</xref>; <xref ref-type="bibr" rid="CIT341">Rod+19</xref>; <xref ref-type="bibr" rid="CIT339">RG22</xref>]. With this tileable input guidance, we can use our autoencoder <inline-formula><mml:math id="M327" display='block'><mml:mi mathvariant="script">A</mml:mi></mml:math></inline-formula> to propagate the neural texture <inline-formula><mml:math id="M328" display='block'><mml:mi mathvariant="script">T</mml:mi></mml:math></inline-formula> , effectively generating tileable BTFs, as we show in <xref ref-type="fig" rid="fig-4.3">Figure 4.3</xref>. This propagation algorithm can leverage state-of-the-art algorithms for tileable texture synthesis without any modification of our material model or training framework. Tileable BTFs were not achievable with previous approximations and this simple pipeline has the potential of enabling novel applications of this type of material representation in rendering scenarios.</p>
</sec>
<sec id="c4-s6.s2">
<label>4.6.2</label>
<title>Structural Material Edition</title>
<p>Besides propagating BTF measurements to larger portions of the material, NeuBTF allows for generating novel materials using structural editions. Given a trained NeuBTF and a guidance image representing some particular target structure, we can propagate the neural texture to this guidance image, generating high-quality neural materials which preserve the structure of the guidance image and the reflectance properties of the trained neural material. This pipeline allows for easily generating multiple different neural materials without the need for retraining. As we show in <xref ref-type="fig" rid="fig-4.8">Figure 4.8</xref>, this propagation method works for many types of input guidances, including vector black and white images, procedurally generated textures, or real photographs of materials and textures. We show results on acquired BTFs and from synthetic BTFs, rendered <target target-type="page" id="pges_114"/>from scanned and manually generated SVBRDFs. As shown, our propagation frameworks provides high-quality material editions, even for very challenging cases, like the <italic>circles pattern</italic>.</p>
<fig id="fig-4.8">
<label>Figure 4.8:</label>
<caption><title>Examples of structural editions allowed by our method on a variety of materials. On the leftmost column, we show <italic>guidance images</italic>, which represent structures into which we transfer the neural BTF measurements illustrated on the top row. As shown, our method can effectively propagate BTF measurements into many material and structure types, using as guidances either synthetic or real images. The first three materials (<italic>leather11, carpet07, fabric01</italic>) are taken from <italic>UBO 2014</italic> [<xref ref-type="bibr" rid="CIT327">Rei+14</xref>], the <italic>linen</italic> material is rendered from a captured SVBRDF, while the <italic>ceramic</italic> material is rendered from an artistic material taken from <italic>Substance3D</italic>. We show renders generated using <inline-formula><mml:math id="M330" display='block'><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo></mml:math></inline-formula>.</title></caption>
<alt-text>Examples of structural editions allowed by our method on a variety of materials. On the leftmost column, we show guidance images, which represent structures into which we transfer the neural BTF measurements illustrated on the top row. As shown, our method can effectively propagate BTF measurements into many material and structure types, using as guidances either synthetic or real images. The first three materials (leather11, carpet07, fabric01) are taken from UBO 2014 [Rei+14], the linen material is rendered from a captured SVBRDF, while the ceramic material is rendered from an artistic material taken from Substance3D. We show renders generated using &#x03B8;v=&#x03B8;l=0..</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.8.jpg"/>
</fig>
</sec>
<sec id="c4-s6.s3">
<label>4.6.3</label>
<title>Multi Resolution Neural Materials</title>
<p>Another useful application enabled by our method is the generation of materials at different resolutions. Unlike previous work [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>], which explicitly optimizes a pyramid of levels of detail during training, we can generate materials at any resolution at test time without introducing any additional complexities to our material representation. Because we train our models using random rescales as a data augmentation policy, they are equivariant to rescales of its input guidance images <inline-formula><mml:math id="M329" display='block'><mml:mi>G</mml:mi><mml:mo>&#x2193;:</mml:mo><mml:mi mathvariant="script">R</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">A</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>&#x2193;=</mml:mo><mml:mi mathvariant="script">R</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">A</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">G</mml:mi><mml:mo stretchy="false">&#x2193;</mml:mo><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:math></inline-formula>. As such, we can generate any continuous resolution for a particular BTF by downsampling the guidance image to the target resolution and propagating its neural texture, as we illustrate in <xref ref-type="fig" rid="fig-4.9">Figure 4.9</xref>. Note that this algorithm only guarantees accurate results for the rescaling ranges that we use during data augmentation.</p>
<fig id="fig-4.9">
<label>Figure 4.9:</label>
<caption><title>Our method naturally enables for the generation of different resolutions for the neural materials. We show the input to the autoencoder (top row), the rendered material at <inline-formula><mml:math id="M331" display='block'><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>75</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>60</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> (middle row), and the ground truth image at those positions (bottom), at different resolutions (columns). We achieve this by downsampling the guidance image <inline-formula><mml:math id="M332" display='block'><mml:mi mathvariant="script">G</mml:mi></mml:math></inline-formula> fed into the autoencoder module, which returns an accurately down-sampled latent space, thanks to our data augmentation policy applied during training. The rightmost column lies beyond the ranges in which we train the model, however, the results are still somewhat plausible.</title></caption>
<alt-text>Our method naturally enables for the generation of different resolutions for the neural materials. We show the input to the autoencoder (top row), the rendered material at &#x03B8;c=75,&#x03B8;l=60,&#x03D5;c=&#x03D5;l=0 (middle row), and the ground truth image at those posi-tions (bottom), at different resolutions (columns). We achieve this by downsampling the guidance image G fed into the autoencoder module, which returns an accurately down-sampled latent space, thanks to our data augmentation policy applied during training. The rightmost column lies beyond the ranges in which we train the model, however, the results are still somewhat plausible.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.9.jpg"/>
</fig>
</sec>
</sec>
<sec id="c4-s7">
<label>4.7</label>
<title>Conclusions</title>
<p>We have presented a learning based representation for material reflectance which provides efficient encoding and powerful propagation capabilities. Our method introduces input conditioning into neural BTF representations. This allows for multiple applications which were not possible with previous neural models, including BTF extrapolation, tiling, and novel material synthesis through structure propagation. Our method builds upon recent work on neural fields, network design and data augmentation, showing competitive compression capabilities with previous work on neural BTF representation. Through multiple analyses, we have shown the capabilities of our method on a variety of materials with different reflectance properties, including anisotropy or specularity, as well as effectively handling either synthetic or acquired BTFs.</p>
<p>Our method can be extended in several ways. The most immediate extension is to allow for materials with strong parallax effects due to displacement mapping or curvature, as in [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>; <xref ref-type="bibr" rid="CIT221">Kuz+22</xref>]. Our representation is limited to opaque materials. Extending them to handle translucent or holed surfaces would increase their realism in materials like thin fabrics or meshes. Further, our method could be extended to allow for hyperspectral BTF data [<xref ref-type="bibr" rid="CIT346">RSK10</xref>], but captured data is scarce. Besides, recent work on neural BRDF representations [<xref ref-type="bibr" rid="CIT380">Szt+21</xref>; <xref ref-type="bibr" rid="CIT315">Rai+22</xref>; <xref ref-type="bibr" rid="CIT088">Fan+22</xref>] and generative models [<xref ref-type="bibr" rid="CIT284">M&#x00FC;l+19</xref>] suggests a promising research direction: Learning to sample from neural BTFs, using invertible neural networks. While these may introduce challenging complexities to the models, they could provide efficient representations for Monte Carlo <target target-type="page" id="pges_115"/>rendering using importance sampling. Further, building upon recent work on SVBRDF capture [<xref ref-type="bibr" rid="CIT337">Rod+23a</xref>; <xref ref-type="bibr" rid="CIT454">Zho+22</xref>], BRDF sampling [<xref ref-type="bibr" rid="CIT295">NJR15</xref>] and BTF compression [<xref ref-type="bibr" rid="CIT316">Rai+20</xref>], it could be possible to learn a prior over neural BTFs with a generative model. This should help in capturing more efficiently the data needed for generating these assets, as well as generating new materials and interpolating between them. Finally, editing semantic and reflectance properties in neural fields is an active area of research [<xref ref-type="bibr" rid="CIT426">WTX22</xref>; <xref ref-type="bibr" rid="CIT251">Liu+21</xref>; <xref ref-type="bibr" rid="CIT407">Wan+22a</xref>; <xref ref-type="bibr" rid="CIT435">Ye+22</xref>; <xref ref-type="bibr" rid="CIT135">Haq+23</xref>]. While our method introduces structural edition into neural BTF representations, it is not capable of editing particular semantic properties, such as albedo or specularity. Extending our edition capabilities to more fine-grained parameters is an interesting research avenue. We hope our method inspires future research on neural material representations.<target target-type="page" id="pges_116"/><target target-type="page" id="pges_117"/></p>
</sec>
<sec id="c4-sA">
<label><target target-type="page" id="pges_118"/>4.A</label>
<title>Model Implementation Details</title>
<fig id="fig-4.10">
<label>Figure 4.10:</label>
<caption><title>A full diagram of our U-Net autoencoder <inline-formula><mml:math id="M333" display='block'><mml:mi mathvariant="script">A</mml:mi></mml:math></inline-formula>. We use a CBAM [<xref ref-type="bibr" rid="CIT425">Woo+18</xref>] module on the bottleneck, Layer Normalization [<xref ref-type="bibr" rid="CIT018">BKH16</xref>], and residual connections in every convolutional block. Inspired by ConvNext Blocks [<xref ref-type="bibr" rid="CIT253">Liu+22b</xref>], we use 5&#x00D7;5 depthwise convolutions followed by 1&#x00D7;1 linear layers, with GeLU non-linearities [<xref ref-type="bibr" rid="CIT147">HG16</xref>]. We use 14 latent dimensions for our neural materials, following [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>]. We use <italic>circular padding</italic> throughout the model and initialize its weights using <italic>orthogonal initialization</italic> [<xref ref-type="bibr" rid="CIT013">ACB19</xref>], which helps avoid exploding gradients. In red, we show the input/output dimensions of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, non-linearities.</title></caption>
<alt-text>A full diagram of our U-Net autoencoder A. We use a CBAM [Woo+18] module on the bottleneck, Layer Normalization [BKH16], and residual connections in every convolu-tional block. Inspired by ConvNext Blocks [Liu+22b], we use 5&#x00D7;5 depthwise convolutions followed by 1&#x00D7;1 linear layers, with GeLU non-linearities [HG16]. We use 14 latent dimen-sions for our neural materials, following [Kuz+21]. We use circular padding throughout the model and initialize its weights using orthogonal initialization [ACB19], which helps avoid exploding gradients. In red, we show the input/output dimensions of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, non-linearities.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.10.jpg"/>
</fig>
<fig id="fig-4.11">
<label><target target-type="page" id="pges_119"/>Figure 4.11:</label>
<caption><title>A full diagram of our sinusoidal renderer <inline-formula><mml:math id="M334" display='block'><mml:mi mathvariant="script">R</mml:mi></mml:math></inline-formula>. We build upon a SIREN architecture [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>], which takes as input the latent vector, queried from the input texture, as well as target camera and light angles to re-render. We introduce Layer Normalization after each linear layer, which we implement using 1 &#x00D7; 1 convolutions to allow for 2-D inputs during training. The output of the model is clamped to avoid negative output radiance values, using a simple ReLU non-linearity [<xref ref-type="bibr" rid="CIT290">NH10</xref>], following [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>].</title></caption>
<alt-text>A full diagram of our sinusoidal renderer R. We build upon a SIREN architecture [Sit+20], which takes as input the latent vector, queried from the input texture, as well as target camera and light angles to re-render. We introduce Layer Normalization after each linear layer, which we implement using 1 &#x00D7; 1 convolutions to allow for 2-D inputs during training. The output of the model is clamped to avoid negative output radiance values, using a simple ReLU non-linearity [NH10], following [Kuz+21].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-4.11.jpg"/>
</fig>
</sec>
</body>
</book-part>
<book-part id="c5" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_120"/><target target-type="page" id="pges_121"/>Chapter 5.</label>
<title>Uncertainty-Aware Neural Networks for High Resolution Digitization</title>
</title-group>
</book-part-meta>
<body>
<fig id="fig-5.1">
<label>Figure 5.1:</label>
<caption><title>From a single image captured with a flatbed scanner <italic>UMat</italic> estimates high-resolution SVBRDFs. Leveraging an attention-guided generative model, the estimations are sharp and artifact-free, and can be used in photorealistic rendering. We introduce a novel uncertainty quantification algorithm for material digitization, which correlates with error at test time and can be used for dataset creation.</title></caption>
<alt-text>From a single image captured with a flatbed scanner UMat estimates high-resolution SVBRDFs. Leveraging an attention-guided generative model, the estimations are sharp and artifact-free, and can be used in photorealistic rendering. We introduce a novel uncertainty quantification algorithm for material digitization, which correlates with error at test time and can be used for dataset creation.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.1.jpg"/>
</fig>
<p>In this chapter, we introduce a learning-based method to recover normals, specularity, and roughness from a single diffuse image of a material, using microgeometry appearance as our primary cue. Previous methods that work on single images tend to produce over-smooth outputs with artifacts, operate at limited resolution, or train one model per class with little room for generalization. In contrast, in this work, we propose a novel capture approach that leverages a generative network with attention and a U-Net discriminator, which shows outstanding performance integrating global information at reduced computational complexity. We showcase the performance of our method with a real dataset of digitized textile materials and show that a commodity flatbed scanner can produce the type of diffuse illumination required as input to our method. Additionally, because the problem might be ill-posed &#x015B;more than a single diffuse image might be needed to disambiguate the specular reflection&#x015B; or because the training dataset is not representative enough of the real distribution, we propose a novel framework to quantify the model&#x2019;s confidence about its predictions at test time. Our method is the first one to deal with the problem of modeling uncertainty in material digitization, increasing the trustworthiness of the process and enabling more intelligent strategies for dataset creation, as we demonstrate with an active learning experiment. A version of the model presented in this chapter is part of <uri xlink:href="http://www.Textura.ai">Textura.ai</uri>. The contributions in this chapter led to the following publication<xref ref-type="fn" rid="c5-fn1"><sup>1</sup></xref>:</p>
<fig id="fig-5">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.jpg"/>
</fig>
<disp-quote>
<p><target target-type="page" id="pges_122"/>&#x201C;UMat: Uncertainty-Aware Single Image High Resolution Material Capture&#x201D;</p>
<p>Carlos Rodriguez-Pardo, Henar Dominguez, David Pascual, Elena Garces</p>
<p><italic>Proc. of Computer Vision and Pattern Recognition (CVPR)</italic> (2023)</p>
</disp-quote>
<fig id="fig-5.2">
<label>Figure 5.2:</label>
<caption><title>Our method digitizes a material taking as input a single scanned image. Further, it returns a pixel-wise metric of uncertainty &#x03C3;<sub>BRDF</sub>, computed at test time through probabilistic sampling, proven useful for active learning. In the plot we compare the average deviations of the radiance of different renders in the blue crop w.r.t the ground truth (GT) of: 1) the distribution of the probabilistic samples of a model trained with 100% of the data; 2) the deterministic output of that model; 3) the output of a model trained using 40% of the training dataset, sampled by active learning guided by &#x03C3;<sub>BRDF</sub> and; 4) a model trained using 40% of the training dataset, randomly sampled. The material at the bottom, for which the model shows a higher uncertainty, generates more varied renders and differs most from the ground truth.</title></caption>
<alt-text>Our method digitizes a material taking as input a single scanned image. Further, it returns a pixel-wise metric of uncertainty &#x03C3;BRDF, computed at test time through probabilistic sampling, proven useful for active learning. In the plot we compare the average deviations of the radiance of different renders in the blue crop w.r.t the ground truth (GT) of: 1) the distribution of the probabilistic samples of a model trained with 100% of the data; 2) the deterministic output of that model; 3) the output of a model trained using 40% of the training dataset, sampled by active learning guided by &#x03C3;BRDF and; 4) a model trained using 40% of the training dataset, randomly sampled. The material at the bottom, for which the model shows a higher uncertainty, generates more varied renders and differs most from the ground truth.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.2.jpg"/>
</fig>
<sec id="c5-s1">
<label><target target-type="page" id="pges_123"/>5.1</label>
<title>Introduction</title>
<p>Virtual design, online marketplaces, product lifecycle workflows, AR/VR, videogames, . . . , all require lifelike digital representations of real-world materials (<italic>i.e.</italic>, digital twins). Acquiring these digital copies is typically a cumbersome and slow process that requires expensive machines and several manual steps, creating roadblocks for scalability, repeatability, and consistency. Among the many industries requiring digital twins of materials, the fashion industry is in a critical position; facing the demand to digitize hundreds of samples of textiles in short periods, which cannot be achieved with the current technology.</p>
<p>In this context, casual capture systems for optical digitization provide a promising path for scalability. These systems leverage handheld devices (such as smartphones), one or more different illuminations, and learning-based priors to estimate the material&#x2019;s diffuse and specular reflection lobes. However, existing approaches present several drawbacks that</p>
<p>make them unsuitable for practical digitization workflows. Generative solutions [<xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT400">VPS21</xref>] typically produce artifacts and are generally not tileable. Despite recent attempts to improve tileability and controllability [<xref ref-type="bibr" rid="CIT454">Zho+22</xref>], these solutions are slow to train and evaluate (requiring online optimization iterations), are limited in resolution, and present challenges for generalization (requiring one model per material class). Further, the fact that these methods build on perceptual losses &#x015B;not pixel losses&#x015B; to compare the input photo with the generated material entails extra difficulties when it comes to guaranteeing the repeatability and consistency required for building a digital inventory (i.e., color swatches, prints, or other variations). On the other hand, methods that build on differentiable node graphs [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>] overcome the tileability and resolution limitations, yet, they share the problems derived from using perceptual losses and category-specific training.</p>
<p>In this chapter, we present <italic>UMat</italic>, a practical, scalable, and reliable approach to digitizing the optical appearance of textile material samples using SVBRDFs. Commonly used SVBRDFs typically contain two reflection terms: a diffuse term, parameterized by an albedo image, and a specular one, parameterized by normals, specularity, and roughness. Prior work typically estimates both <target target-type="page" id="pges_124"/>components, which becomes a very challenging problem, obtaining over-smooth outputs, and being prone to artifacts [<xref ref-type="bibr" rid="CIT065">Des+18</xref>; <xref ref-type="bibr" rid="CIT125">Guo+21</xref>; <xref ref-type="bibr" rid="CIT456">ZK21</xref>]. Instead, in this chapter, we demonstrate that it is possible to provide accurate digitizations of materials leveraging as input a single diffuse image that acts as albedo and estimating the specular components using a neural network.</p>
<p>Our key observation is to realize that most of the appearance variability of textile materials is due to its microgeometry and that a commodity flatbed scanner can approximate the type of diffuse illumination that we require for the majority of textile materials (see <xref ref-type="fig" rid="fig-5.3">Figure 5.3</xref>).</p>
<fig id="fig-5.3">
<label>Figure 5.3:</label>
<caption><title>Scanner images vs fitted albedos.</title></caption>
<alt-text>Scanner images vs fitted albedos.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.3.jpg"/>
</fig>
<p>Nevertheless, single-image material estimation is still an ill-posed problem in our setting, as reflectance properties may not be directly observable from a single diffuse image. To account for these non-directly observable properties, we propose a novel way to measure the model&#x2019;s confidence about its prediction at test time. Leveraging Monte Carlo (MC) Dropout [<xref ref-type="bibr" rid="CIT097">GG16</xref>], we propose an uncertainty metric computed as the variance of sampling and evaluating multiple estimations for a single input in a render space. We show that this confidence directly correlates with the accuracy of the digitization, which helps identify ambiguous inputs, out-of-distribution samples, or under-represented classes. Besides increasing the trustworthiness of the capture process, our confidence quantification enables smarter strategies for dataset creation, as we demonstrate with an active learning experiment.</p>
<p>We pose the estimation as an Image-to-Image Translation problem (I2IT) that directly re-gresses roughness, specular, and normals, from a single input image. Under the hood, our novel residual architecture has a single encoder enhanced with lightweight attention modules [<xref ref-type="bibr" rid="CIT415">Wan+20b</xref>; <xref ref-type="bibr" rid="CIT268">MR21</xref>] for improving global consistency and reducing artifacts, specialized decoders for each target reflectance map, and a U-Net discriminator [<xref ref-type="bibr" rid="CIT354">SSK20</xref>], which enhances generalization.</p>
<p>In summary, we present the following contributions:</p>
<list list-type="bullet">
<list-item><p>A novel material capture system which leverages the diffuse illumination provided by flatbed scanners for high-resolution, scalable, and reliable digitizations.</p></list-item>
<list-item><p>An attention-enhanced GAN model and training procedure designed for maximizing accuracy and sharpness, and removing undesired artifacts.</p></list-item>
<list-item><p>A generic uncertainty quantification framework for material capture algorithms which correlates with prediction error on a render space.</p></list-item>
<list-item><p>An exhaustive evaluation method for measuring the accuracy and quality of the model estimations.</p></list-item>
</list>
</sec>
<sec id="c5-s2">
<label><target target-type="page" id="pges_125"/>5.2</label>
<title>Related Work</title>
<sec id="c5-s2.s1">
<label>5.2.1</label>
<title>Lightweight Material Capture</title>
<p><bold>Single Image SVBRDF Capture</bold> Capturing accurate SVBRDFs from a single image is a challenging problem that requires predicting the photometric response of a material given only a sample of it. The most common approximation is to use a flash-lit front planar image captured with a smartphone. Extending neural style transfer [<xref ref-type="bibr" rid="CIT108">GEB15b</xref>] to material capture, Aittala <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT003">AAL16</xref>] leverage pre-trained CNNs and texture priors for smartphone material acquisition. Relatedly, Henzler <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>] use style losses for training generative models for BRDF synthesis. These approaches are optimized for synthesis, which allow for seamlessly tileable outputs, and do not require supervised training. However, they work best for stochastic materials, limiting their scope.</p>
<p>Leveraging datasets of labeled materials, different methods have trained autoencoders for SVBRDF capture. Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT235">Li+17a</xref>] reconstruct spatially-varying albedo and normals using a U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>], and homogeneous specular albedo and roughness using a CNN regressor. This work was extended through Self-Augmented CNNs in [<xref ref-type="bibr" rid="CIT437">Ye+18</xref>]. By leveraging Conditional Random Fields (CRFs), a material classifier and one decoder per map, Li <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT242">LSC18</xref>] reconstruct spatially-varying albedo, roughness, and normals. Deschaintre <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT065">Des+18</xref>] propose a modified U-Net, synthetic datasets, and a render loss. Cascaded models [<xref ref-type="bibr" rid="CIT243">Li+18</xref>; <xref ref-type="bibr" rid="CIT351">SC20</xref>]; and deep latent spaces optimized using inverse rendering [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>] have also shown success for this problem. Recently, Generative Adversarial Networks (GANs) have shown improved capabilities compared to more naive losses. These require a discriminator, which can be trained on renders [<xref ref-type="bibr" rid="CIT424">Wen+22</xref>], SVBRDF maps [<xref ref-type="bibr" rid="CIT125">Guo+21</xref>; <xref ref-type="bibr" rid="CIT400">VPS21</xref>], or both [<xref ref-type="bibr" rid="CIT456">ZK21</xref>].</p>
<p>Our approach differs from previous work in several factors. Importantly, we use flatbed scanners instead of smartphones for material capture. While they limit the materials which can be captured, they provide adequate illumination for easier digitizations, and a higher level of resolution and detail. We hypothesize that material specularity can be estimated accurately by leveraging its microgeometry. From this assumption, we build a GAN which, in contrast with previous work, leverages state-of-the-art attention mechanisms and discriminator design for obtaining a more holistic understanding of its inputs. Further, we train exclusively on real data, and propose a more comprehensive evaluation.</p>
<p><bold>Multiple Image SVBRDF Capture</bold> A different corpus of work relies on multiple images of the material. By capturing more evidence of the material photometric response, they provide more accurate SVBRDFs. This hinders scalability, as they require a larger capture and calibration effort. The simplest approach is to require two samples, a diffuse and a flash-lit image [<xref ref-type="bibr" rid="CIT005">AWL15</xref>; <xref ref-type="bibr" rid="CIT036">Bos+20</xref>]. More <target target-type="page" id="pges_126"/>flexible alternatives allow for more samples, combined with a learned prior [<xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT066">Des+19</xref>]; or Monte Carlo rendering [<xref ref-type="bibr" rid="CIT256">Lua+21</xref>]. A different approach is to capture the material at a high resolution, and transfer those details to larger samples of the same material [<xref ref-type="bibr" rid="CIT067">DDB20</xref>; <xref ref-type="bibr" rid="CIT338">RG21</xref>]. Finally, videos for material acquisition have also shown to be promising capture systems [<xref ref-type="bibr" rid="CIT436">Ye+21</xref>].</p>
<p><bold>Procedural Graphs</bold> Procedural graphs have also been used for material generation. Instead of relying on material priors, these approaches work by optimizing a material graph through a differentiable pipeline. These provide interesting capabilities, such as easy edition or tiling, but are limited by the expressiveness of the procedural model. These have been explored for general SVBRDF estimation [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>; <xref ref-type="bibr" rid="CIT165">Hu+22c</xref>; <xref ref-type="bibr" rid="CIT126">Guo+20a</xref>], or for high-quality woven fabric digitizations [<xref ref-type="bibr" rid="CIT184">Jin+22</xref>].</p>
</sec>
<sec id="c5-s2.s2">
<label>5.2.2</label>
<title>Uncertainty Quantification in Deep Learning</title>
<p>Measuring the confidence of deep learning models is an active research area [<xref ref-type="bibr" rid="CIT203">KG17</xref>] with multiple applications, including safety-critical problems, like self-driving [<xref ref-type="bibr" rid="CIT170">Hua+19</xref>] or medicine [<xref ref-type="bibr" rid="CIT219">Kur+22</xref>]; and active dataset creation [<xref ref-type="bibr" rid="CIT232">Lei+21</xref>; <xref ref-type="bibr" rid="CIT370">Sol+21</xref>]. In computer vision, uncertainty quantification</p>
<p>has focused on image classification [<xref ref-type="bibr" rid="CIT358">SKK18</xref>; <xref ref-type="bibr" rid="CIT232">Lei+21</xref>], segmentation [<xref ref-type="bibr" rid="CIT213">Kry+21</xref>], and depth regression [<xref ref-type="bibr" rid="CIT160">Hu+22a</xref>; <xref ref-type="bibr" rid="CIT009">Ami+20</xref>]. To overcome the computational intractability of Bayesian Neural Networks, different approximations have been proposed, including MC Dropout [<xref ref-type="bibr" rid="CIT097">GG16</xref>], Deep Ensembles [<xref ref-type="bibr" rid="CIT227">LPB17</xref>] or Variational Inference [<xref ref-type="bibr" rid="CIT273">MNG17</xref>]. Orthogonal alternatives exist, including Evidential Deep Learning [<xref ref-type="bibr" rid="CIT009">Ami+20</xref>; <xref ref-type="bibr" rid="CIT358">SKK18</xref>] or frequentist approaches like Constrained Ordinal Regression [<xref ref-type="bibr" rid="CIT160">Hu+22a</xref>]. We refer the reader to recent surveys [<xref ref-type="bibr" rid="CIT187">Jos+22</xref>; <xref ref-type="bibr" rid="CIT113">Gaw+21</xref>] for more comprehensive reviews.</p>
<p><target target-type="page" id="pges_127"/>Quantifying uncertainties allows communicating the end user that the model predictions may be inaccurate, suggesting alternative pathways; as well as cheaper dataset creation through <italic>active learning</italic>. Bayesian material parameter estimation has been proposed for procedural frameworks [<xref ref-type="bibr" rid="CIT126">Guo+20a</xref>], but, to the best of our knowledge, it has not been explored for SVBRDF estimation. We aim to propose an efficient uncertainty quantification framework for deep SVBRDF capture methods which accounts for material perception and is agnostic to the material model.</p>
</sec>
</sec>
<sec id="c5-s3">
<label>5.3</label>
<title>Method</title>
<p>Our method starts from an input image X of a material taken under diffuse lighting and outputs the parameters of the specular lobe of the SVBRDF, <italic>i.e.</italic>, <inline-formula><mml:math id="M335" display='block'><mml:mi mathvariant="bold">M</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi mathvariant="script">M</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>3</mml:mn></mml:msubsup></mml:math></inline-formula> corresponding to the material roughness, specularity, and normals. We illustrate this process in <xref ref-type="fig" rid="fig-5.4">Figure 5.4</xref>.</p>
<fig id="fig-5.4">
<label>Figure 5.4:</label>
<caption><title>Overview of <italic>UMat</italic>. We propose an attention-guided generator trained with style, frequency, pixel-wise, and adversarial losses. In green, we show the components that include any form of attention mechanism. The supplementary material contains the detailed architectures. On the right, we show two applications of our method: First, thanks to our test-time uncertainty quantification, we can provide a measure of the reliability of the estimation. Second, the maps that <italic>UMat</italic> produces can be used by any render engine.</title></caption>
<alt-text>Overview of UMat. We propose an attention-guided generator trained with style, frequency, pixel-wise, and adversarial losses. In green, we show the components that include any form of attention mechanism. The supplementary material contains the detailed architectures. On the right, we show two applications of our method: First, thanks to our test-time uncertainty quantification, we can provide a measure of the reliability of the estimation. Second, the maps that UMat produces can be used by any render engine.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.4.jpg"/>
</fig>
<p>Following previous work [<xref ref-type="bibr" rid="CIT192">KG13</xref>; <xref ref-type="bibr" rid="CIT286">Mun+22</xref>], we use the physically-based material model from Disney [<xref ref-type="bibr" rid="CIT038">BS12</xref>], which aggregates a diffuse term with an isotropic, microfacet specular GGX lobe <inline-formula><mml:math id="M336" display='block'><mml:mi>s</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="CIT405">Wal+07</xref>], such that, <inline-formula><mml:math id="M337" display='block'><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">M</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">X</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mi mathvariant="normal">X</mml:mi><mml:mi>&#x03C0;</mml:mi></mml:mfrac><mml:mo>+</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">M</mml:mi><mml:mo>)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the shading model for a light <italic>l</italic> and view position <italic>v</italic>. We formulate this estimation as an Image-to-Image Translation problem (I2IT). We train a U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] generator <inline-formula><mml:math id="M338" display='block'><mml:mi>G</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mover><mml:mi>M</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> within a GAN framework, extending the adversarial loss with pixel-wise, style, and frequency losses. We design the generator to maximize accuracy, sharpness, and robustness. To do so, we train a single encoder, enhanced with self-attention layers and a transformer, and use one decoder per map. <xref ref-type="sec" rid="c5-s3.s1">Sections 5.3.1</xref>, <xref ref-type="sec" rid="c5-s3.s2">5.3.2</xref>, and <xref ref-type="sec" rid="c5-s3.s3">5.3.3</xref> present the network design, the loss and data augmentation choices, respectively. Implementation details are provided on the supplementary material (<xref ref-type="sec" rid="c5.sA">Section 5.A</xref>).</p>
<p>Using as input a single image taken under diffuse lighting presents extra challenges when estimating the SVBRDF; we lack the extra cues provided by more complex illumination patterns (e.g. flash lighting). Therefore, in <xref ref-type="sec" rid="c5-s4">Section 5.4</xref>, we explain how to compensate for this potential ambiguity by introducing an uncertainty metric that can be computed at test time. <xref ref-type="sec" rid="c5-s5">Section 5.5</xref> presents the evaluation, which includes the description of our dataset and metrics (<xref ref-type="sec" rid="c5-s5.s1">Sections 5.5.1</xref> and <xref ref-type="sec" rid="c5-s5.s2">5.5.2</xref>), an ablation study that validates the design (<xref ref-type="sec" rid="c5-s6.s1">5.6.1</xref>), qualitative and quantitative results (<xref ref-type="sec" rid="c5-s6.s2">5.6.2</xref> and <xref ref-type="sec" rid="c5-s6.s3">5.6.3</xref>), an application of our uncertainty metric in an active learning setting (<xref ref-type="sec" rid="c5-s6.s4">5.6.4</xref>), and comparisons with previous work (<xref ref-type="sec" rid="c5-s6.s5">5.6.5</xref>).</p>
<sec id="c5-s3.s1">
<label>5.3.1</label>
<title>Network Design</title>
<p>We use a U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] with residual connections [<xref ref-type="bibr" rid="CIT139">He+16</xref>; <xref ref-type="bibr" rid="CIT069">Dia+20</xref>; DLG21] in all of our convolutional blocks, an individual decoder per map [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>; <xref ref-type="bibr" rid="CIT339">RG22</xref>; <target target-type="page" id="pges_128"/><xref ref-type="bibr" rid="CIT068">DLG21</xref>; <xref ref-type="bibr" rid="CIT456">ZK21</xref>], and group normalization [<xref ref-type="bibr" rid="CIT427">WH18</xref>]. To each decoder, we add a pixel-wise dropout-regularized MLP [<xref ref-type="bibr" rid="CIT373">Sri+14</xref>], aimed at increasing the accuracy of the predictions and allowing us to measure uncertainty at test-time.</p>
<p>While this multi-decoder residual U-Net is relatively accurate, it is limited by its receptive field, as is common on fully-convolutional architectures. Previous work [<xref ref-type="bibr" rid="CIT065">Des+18</xref>] proposed the use of a <italic>global track</italic> for fusing spatially distant information. We instead draw inspiration from recent advances in attention and diffusion models [<xref ref-type="bibr" rid="CIT349">Sah+22</xref>; <xref ref-type="bibr" rid="CIT157">HJA20</xref>; <xref ref-type="bibr" rid="CIT343">Rom+22</xref>], and add a self-attention module with linear complexity [<xref ref-type="bibr" rid="CIT415">Wan+20b</xref>] to the output of every convolutional block in the encoder. Finally, we add a lightweight MobileViT [<xref ref-type="bibr" rid="CIT268">MR21</xref>] to the bottleneck to provide the model with a global understanding of its input. By performing the most complex computations at the encoder, we provide the specialized decoders with dense inputs which are computed only once.</p>
</sec>
<sec id="c5-s3.s2">
<label>5.3.2</label>
<title>Loss Function</title>
<p><disp-formula id="Eq005-1"><label>(5.1)</label> <mml:math id="M339" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mtext>pixel</mml:mtext><mml:mi>i</mml:mi></mml:msub></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>freq</mml:mtext></mml:msub><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>freq</mml:mtext></mml:msub></mml:math></disp-formula></p>
<p><inline-formula><mml:math id="M340" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mtext>pixel</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the <inline-formula><mml:math id="M341" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> norm weighted per map, <italic>&#x03BB;<sub>i</sub></italic>. <inline-formula><mml:math id="M342" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> produces sharper results than higher-order alternatives, such as <inline-formula><mml:math id="M343" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>. We introduce an adversarial loss to handle the intrinsic ambiguity of ill-posed problems [<xref ref-type="bibr" rid="CIT387">Tex+20b</xref>; <xref ref-type="bibr" rid="CIT176">Iso+17</xref>; <xref ref-type="bibr" rid="CIT104">Gar+22</xref>; <xref ref-type="bibr" rid="CIT400">VPS21</xref>]. In our case, the choice of the discriminator is a critical design decision.</p>
<p>Recent work [<xref ref-type="bibr" rid="CIT354">SSK20</xref>] proposed U-Net architectures for discriminators, which allows for better learn both low and high-level features, and introduces further regularization. These result in more conservative albeit less diverse generations [<xref ref-type="bibr" rid="CIT133">Han+22a</xref>]. We use a U-Net discriminator [<xref ref-type="bibr" rid="CIT354">SSK20</xref>] with attention [<xref ref-type="bibr" rid="CIT425">Woo+18</xref>], which outputs two estimations: a scalar output <inline-formula><mml:math id="M344" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, provided by its encoder, and a 2D estimation <inline-formula><mml:math id="M345" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, provided by its decoder. <inline-formula><mml:math id="M346" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> provides a global estimate of the quality of the stack <inline-formula><mml:math id="M347" display='block'><mml:mtext>M</mml:mtext></mml:math></inline-formula>, while <inline-formula><mml:math id="M348" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> gives pi estimations. As in [<xref ref-type="bibr" rid="CIT354">SSK20</xref>], we add a <italic>regularization</italic> term <inline-formula><mml:math id="M349" display='block'><mml:msubsup><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mtext>dec</mml:mtext></mml:msub><mml:mtext>cons</mml:mtext></mml:msubsup></mml:math></inline-formula>, and leverage <italic>cut-mix</italic> as for discriminator data-augmentation. Our discriminator and adversarial losses are:</p>
<p><disp-formula id="Eq005-2"><label>(5.2)</label> <mml:math id="M350" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi>D</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mi>D</mml:mi><mml:mtext>enc</mml:mtext></mml:msub></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mi>D</mml:mi><mml:mtext>dec</mml:mtext></mml:msub></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>cons</mml:mtext></mml:msub><mml:msubsup><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mi>D</mml:mi><mml:mtext>dec</mml:mtext></mml:msub><mml:mtext>cons</mml:mtext></mml:msubsup></mml:math></disp-formula></p>
<p><disp-formula id="Eq005-3"><label>(5.3)</label> <mml:math id="M351" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mtext>enc</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="normal">G</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">X</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mtext>dec</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="normal">G</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">X</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula></p>
<p>where the implementation of <inline-formula><mml:math id="M352" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M353" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:msub></mml:math></inline-formula> follows [<xref ref-type="bibr" rid="CIT354">SSK20</xref>].</p>
<p><target target-type="page" id="pges_129"/>We further add two losses to improve the accuracy and sharpness of the results: a frequency loss <inline-formula><mml:math id="M354" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> and a style loss <inline-formula><mml:math id="M355" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mtext>style</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>. <inline-formula><mml:math id="M356" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is estimated by averaging the <italic>focal frequency loss</italic> [<xref ref-type="bibr" rid="CIT183">Jia+21</xref>] computed over each individual channel of M. This is designed to help GANs preserve high-frequency details. Further, as shown in prior work, working in the frequency domain is beneficial when handling textures [<xref ref-type="bibr" rid="CIT004">AWL13</xref>; <xref ref-type="bibr" rid="CIT263">Mar+20</xref>]. Our style loss <inline-formula><mml:math id="M358" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mtext>style</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is inspired by the success of neural losses when dealing with textures [<xref ref-type="bibr" rid="CIT108">GEB15b</xref>; <xref ref-type="bibr" rid="CIT107">GEB15a</xref>]. However, off-the-shelf metrics which are designed for 3-channel images, are not immediately usable in SVBRDF. While it is possible to compute them for each map separetely [<xref ref-type="bibr" rid="CIT339">RG22</xref>], this does not necessarily preserve inter-map consistency. We follow recent work [<xref ref-type="bibr" rid="CIT041">CHB21</xref>] and use LPIPS [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] taking as input a 3-channel image created by randomly sampling three channels from the set of five available channels of M.</p>
</sec>
<sec id="c5-s3.s3">
<label>5.3.3</label>
<title>Data Augmentation</title>
<p>We follow two strategies for data augmentation. First, we perform patch-based training and affine transforms with randomly cropped patches [<xref ref-type="bibr" rid="CIT338">RG21</xref>; <xref ref-type="bibr" rid="CIT387">Tex+20b</xref>; <xref ref-type="bibr" rid="CIT400">VPS21</xref>]. We also apply random rescales for generalization at lower resolutions, and rotations to account for possible misalignments that may arise when capturing the samples. Second, we apply several image transformations to increase model robustness: random intensity changes in HSV space to the inputs, gaussian noise and blurs, and random erasing [<xref ref-type="bibr" rid="CIT452">Zho+20</xref>] for regularization.</p>
</sec>
</sec>
<sec id="c5-s4">
<label>5.4</label>
<title>Uncertainty Quantification</title>
<p>Material acquisition from a single diffuse image is potentially an ill-posed problem, since it assumes that the microgeometry is a sufficient cue to predict the material appearance. While purely learned priors have been proven to work well for inverse problems of single images object shape estimation [<xref ref-type="bibr" rid="CIT246">Lic+21</xref>; <xref ref-type="bibr" rid="CIT256">Lua+21</xref>; <xref ref-type="bibr" rid="CIT173">Hwa+22</xref>; <xref ref-type="bibr" rid="CIT286">Mun+22</xref>; <xref ref-type="bibr" rid="CIT243">Li+18</xref>], they are typically combined <target target-type="page" id="pges_130"/>with render losses that guarantee consistency in the reconstruction. We lack the necessary input to include this kind of supervision, therefore, we propose an alternative approach aimed at quantifying the confidence of the prediction.</p>
<p>This is valuable for several purposes. First, it provides a way of communicating possible inaccuracies to users of these systems when it is not possible to have access to the ground truth reflectance values; and, importantly, it enables efficient dataset creation through active learning, as we show in <xref ref-type="sec" rid="c5-s6.s4">Section 5.6.4</xref>.</p>
<p>We propose an uncertainty quantification mechanism that is applied to individual per-map estimations, and also globally in a render space.</p>
<p>It is possible to measure uncertainty, among other methods, through deep ensembles [<xref ref-type="bibr" rid="CIT227">LPB17</xref>], evidential learning [<xref ref-type="bibr" rid="CIT358">SKK18</xref>; <xref ref-type="bibr" rid="CIT021">BYK21</xref>; <xref ref-type="bibr" rid="CIT408">Wan+22b</xref>], or ordinal regression [<xref ref-type="bibr" rid="CIT160">Hu+22a</xref>]. However, they are costly to train and evaluate, and would imply major changes in our method. Instead, we follow a probabilistic approach called MC Dropout [<xref ref-type="bibr" rid="CIT097">GG16</xref>], with which, for a particular input, we sample a set of predictions by adding randomness to the forward pass of our model. This process has no impact in the regular deterministic evaluation and implies no changes in our model architecture. Specifically, while measuring uncertainty, we randomly deactivate 20% of the neurons of the MLP of our decoders, obtaining a set of outputs <inline-formula><mml:math id="M360" display='block'><mml:mi mathvariant="normal">U</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mover><mml:mi mathvariant="script">M</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x007E;</mml:mo><mml:mi mathvariant="normal">G</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">X</mml:mi><mml:mo>)</mml:mo><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup></mml:math></inline-formula> that we use to compute several metrics.</p>
<p>First, we compute the pixel-wise standard deviation for each invididual map separately, obtaining <inline-formula><mml:math id="M361" display='block'><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mspace/><mml:mspace/><mml:mspace/><mml:mspace/><mml:mspace/><mml:mo>)</mml:mo></mml:mrow><mml:mspace/></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>spec</mml:mi></mml:msub></mml:math></inline-formula>, and &#x03C3;<sub>rough</sub>, for the normals, specular, and roughness maps, respectively. Then, inspired by perceptual metrics for computing BRDF differences [<xref ref-type="bibr" rid="CIT229">Lav+21</xref>], we define our novel perceptually-aware uncertainty metric &#x03C3;<sub>BRDF</sub> as follows:</p>
<p><disp-formula id="Eq005-4"><label>(5.4)</label> <mml:math id="M362" display='block'><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mi>BRDF</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mi>y</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:munder><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:mi>S</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:msqrt><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:munder><mml:mroot><mml:mrow><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">U</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>3</mml:mn></mml:mroot></mml:msqrt><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>where <italic>K</italic> is a 2D grayscale image of a constant neutral grey value, <inline-formula><mml:math id="M363" display='block'><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the pixel-wise standard deviation of the renders obtained for the set of sampled maps U, and <italic>S</italic> is the fixed set of 50 views optimized in [<xref ref-type="bibr" rid="CIT295">NJR15</xref>] for efficient BRDF capture. The log spatially-varying result is integrated across the spatial dimensions <italic>xy</italic> to obtain a single value as output. This equation introduces a perceptual component to our uncertainty metric in two forms: first, by applying cosine weighting to the light position to compensate for light attenuation at grazing angles, and second, by taking the cubic root of these differences to attenuate peak reflectances. <xref ref-type="fig" rid="fig-5.5">Figure 5.5</xref> showcases an <target target-type="page" id="pges_131"/>example of the predicted uncertainties for a rib material with sequins. While the uncertainty is low at the yarns because it is a common material in our dataset, it appears high in the sequins since that effect had not been observed during training.</p>
<fig id="fig-5.5">
<label>Figure 5.5:</label>
<caption><title>Top: input image of a <italic>rib</italic> material with metallic sequins. Bottom: &#x03C3;<sub>BRDF</sub> and per-map uncertainties.</title></caption>
<alt-text>Top: input image of a rib material with metallic sequins. Bottom: &#x03C3;BRDF and per-map uncertainties.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.5.jpg"/>
</fig>
</sec>
<sec id="c5-s5">
<label>5.5</label>
<title>Experimental Results</title>
<sec id="c5-s5.s1">
<label>5.5.1</label>
<title>Dataset</title>
<p>We gathered a novel dataset for training and testing our method. It comprises 2000 textile materials of a variety of families and microstructures that we divided into 14 families: crepe, jacquard, pile, plain, satin, and twill for <italic>wovens</italic>; fleece, french terry, interlock, jersey, milano, pique, and rib for <italic>knits</italic>; and, finally, <italic>leathers</italic>. Each family differs in its construction pattern</p>
<p>(<italic>i.e.,</italic> its microstructure), which directly impacts its optical appearance. In the supplementary material, we show a more detailed analysis of the dataset. For each material, we have an image scanned with the flatbed scanner EPSON V850 whose lighting configuration is close to diffuse (as we show in <xref ref-type="fig" rid="fig-5.3">Figure 5.3</xref>), and its corresponding ground truth specular maps of the SVBRDF (normals, specular, and roughness). To obtain these maps, first, we digitized the material with an optical gonioreflectometer and then, propagated the maps to the scan using map propagation techniques [<xref ref-type="bibr" rid="CIT338">RG21</xref>]. All our images and maps have a resolution of 1000 PPIs, allowing us to leverage the full semantics of the microstructure for inference. We split our dataset into 90-10 for train and test, making sure that every family is equally represented in both splits.</p>
</sec>
<sec id="c5-s5.s2">
<label>5.5.2</label>
<title>Metrics</title>
<p>We quantify individual <italic>per-map accuracy</italic>, <italic>rendered perceptual accuracy</italic>, and <italic>artifacts</italic>. <bold>Per-map accuracy</bold> is computed differently depending on the semantics of the map: the roughness and specular maps are evaluated using Mean Absolute Error (<inline-formula><mml:math id="M364" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>), and the normals are evaluated using the angular distance in vector space (<inline-formula><mml:math id="M365" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mspace/><mml:mspace/><mml:mspace/><mml:mspace/><mml:mspace/><mml:mo>)</mml:mo></mml:mrow><mml:mspace/></mml:mrow></mml:msub></mml:math></inline-formula>). Further, to account for the possibility that the model always returns an accurate average value, resulting in relatively low <inline-formula><mml:math id="M366" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula>, we also measure Pearson correlation <italic>&#x03C1;</italic>.</p>
<p>We evaluate <bold>rendered perceptual accuracy</bold> <inline-formula><mml:math id="M367" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mtext>BRDF</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> between the ground truth stack <inline-formula><mml:math id="M368" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the estimation <inline-formula><mml:math id="M369" display='block'><mml:mover><mml:mi mathvariant="bold">M</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> following existing metrics [<xref ref-type="bibr" rid="CIT229">Lav+21</xref>],</p>
<p><disp-formula id="Eq005-5"><label>(5.5)</label> <mml:math id="M370" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi>BRDF</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mi>y</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:munder><mml:msqrt><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:mi>S</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:munder><mml:mroot><mml:mrow><mml:msup><mml:mi>cos</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold">M</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mover><mml:mi mathvariant="bold">M</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mn>3</mml:mn></mml:mroot></mml:msqrt></mml:math></disp-formula> where the terms are the same as in <xref ref-type="disp-formula" rid="Eq005-4">Equation 5.4</xref>.</p>
<p><target target-type="page" id="pges_132"/>Finally, we found that some architectural improvements in the neural network introduce artifacts in the specular and roughness maps that were not present in the input image, as illustrated in <xref ref-type="fig" rid="fig-5.6">Figure 5.6</xref>. Thus, we provide an <bold>artifacts detection</bold> metric to quantify them. We start by defining a metric of homogeneity for an input image <inline-formula><mml:math id="M371" display='block'><mml:mi mathvariant="script">H</mml:mi></mml:math></inline-formula>(I),</p>
<p><disp-formula id="Eq005-6"><label>(5.6)</label> <mml:math id="M372" display='block'><mml:mi mathvariant="script">H</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">I</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mi>y</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:munder><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mo>{</mml:mo><mml:mo stretchy="false">&#x2191;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">&#x2193;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mo>}</mml:mo></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mo>&#x2225;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mtext>Box</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="normal">I</mml:mi><mml:mo>)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mtext>Box</mml:mtext></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="normal">I</mml:mi><mml:mi>d</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></disp-formula> where I<sup><italic>d</italic></sup> is the image shifted up, down, left, and right by a number of pixels equal to the kernel size of the box filter. Then, we define three metrics that we compute per map: <inline-formula><mml:math id="M373" display='block'><mml:msub><mml:mi>e</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="script">H</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula>, <inline-formula><mml:math id="M374" display='block'><mml:msub><mml:mi>e</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="script">H</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="script">H</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>and</mml:mtext><mml:msub><mml:mi>e</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>I</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">X</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. X is the original input image, and <italic>MI</italic> is the Mutual Information [<xref ref-type="bibr" rid="CIT348">Rus+04</xref>], which helps discern artifacts that appear in a single map from semantic patterns (<italic>e.g.,</italic> plaids, prints). Each map is labeled as having artifacts if the majority of metrics exceed their corresponding thresholds, <inline-formula><mml:math id="M375" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="script">M</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula>. If any of the maps contain artifacts, the entire stack <inline-formula><mml:math id="M376" display='block'><mml:mtext>M</mml:mtext></mml:math></inline-formula> is classified as having artifacts.</p>
<fig id="fig-5.6">
<label>Figure 5.6:</label>
<caption><title>Qualitative results of some configurations of our ablation study. In <bold>red</bold>, we show that the baseline generator architecture trained on the full loss introduces artifacts, which are removed using <bold>attention</bold> on the encoder.</title></caption>
<alt-text>Qualitative results of some configurations of our ablation study. In red, we show that the baseline generator architecture trained on the full loss introduces artifacts, which are removed using attention on the encoder.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.6.jpg"/>
</fig>
</sec>
</sec>
<sec id="c5-s6">
<label>5.6</label>
<title>Evaluation</title>
<sec id="c5-s6.s1">
<label>5.6.1</label>
<title>Ablation Study</title>
<p>In <xref ref-type="fig" rid="fig-5.6">Figure 5.6</xref> and <xref ref-type="table" rid="c5-tab1">Table 5.1</xref>, we present an ablation study to validate our model design. From a baseline U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] trained with a pixel-wise loss and no data augmentation other than rescales, we add different components to the model to improve its generalization. First, we observe that using a PatchGAN [<xref ref-type="bibr" rid="CIT176">Iso+17</xref>; <xref ref-type="bibr" rid="CIT400">VPS21</xref>] discriminator provides a significant increase in accuracy. However, using a similarly-sized U-Net discriminator [<xref ref-type="bibr" rid="CIT354">SSK20</xref>], we achieve better results, particularly when using discriminator regularization. Further, <inline-formula><mml:math id="M377" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mtext>style</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M378" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mrow><mml:mtext>freq</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> yield significant improvements, most notably in the normal map. To the baseline U-Net trained with the full loss, adding residual connections and one decoder per map increases accuracy. However, this setup tends to produce artifacts as it struggles to integrate global information. By adding self-attention to the encoder, we remove the artifacts; and with the MobileViT [<xref ref-type="bibr" rid="CIT268">MR21</xref>], we achieve higher quality results. Our final model contains full data augmentation, which provides small gains on generalization.</p>
<table-wrap id="c5-tab1">
<label>Table 5.1:</label>
<caption><title>Results of our ablation study, across a variety of metrics. <italic>Art.</italic> refers to our artifact detection metric. We use a color code to highlight <bold>best</bold> and <bold>worst</bold> cases.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="left" colspan="2"><p><bold>Configuration</bold></p></th>
<th valign="top" align="center"><p><inline-formula><mml:math id="M379" display='block'><mml:msup><mml:mi>&#x03C1;</mml:mi><mml:mi>S</mml:mi></mml:msup><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></p></th>
<th valign="top" align="center"><p><inline-formula><mml:math id="M380" display='block'><mml:msup><mml:mi>&#x03C1;</mml:mi><mml:mi>R</mml:mi></mml:msup><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></p></th>
<th valign="top" align="center"><p><inline-formula><mml:math id="M381" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mspace/><mml:mspace/><mml:mspace/><mml:mspace/><mml:mspace/><mml:mo>)</mml:mo></mml:mrow><mml:mspace/></mml:mrow></mml:msub><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></p></th>
<th valign="top" align="center"><p><inline-formula><mml:math id="M382" display='block'><mml:msubsup><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn><mml:mi>S</mml:mi></mml:msubsup><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></p></th>
<th valign="top" align="center"><p><inline-formula><mml:math id="M383" display='block'><mml:msubsup><mml:mi mathvariant="script">L</mml:mi><mml:mn>1</mml:mn><mml:mi>R</mml:mi></mml:msubsup><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></p></th>
<th valign="top" align="center"><p><inline-formula><mml:math id="M384" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>BRDF</mml:mtext></mml:msub><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></p></th>
<th valign="top" align="center"><p>Art.&#x2193;</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center"/>
<td valign="top" align="left" colspan="2"><p>Baseline</p></td>
<td valign="top" align="center"><p><bold>0.615</bold></p></td>
<td valign="top" align="center"><p><bold>0.329</bold></p></td>
<td valign="top" align="center"><p><bold>7.570</bold></p></td>
<td valign="top" align="center"><p><bold>0.130</bold></p></td>
<td valign="top" align="center"><p><bold>0.070</bold></p></td>
<td valign="top" align="center"><p><bold>0.325</bold></p></td>
<td valign="top" align="center"><p><bold>0.0</bold></p></td>
</tr>
<tr>
<td valign="top" align="left" rowspan="5"><p><bold>Loss</bold></p></td>
<td valign="top" align="right"><p>Baseline +</p></td>
<td valign="top" align="right"><p>PatchGAN</p></td>
<td valign="top" align="center"><p>0.810</p></td>
<td valign="top" align="center"><p>0.510</p></td>
<td valign="top" align="center"><p>3.729</p></td>
<td valign="top" align="center"><p>0.089</p></td>
<td valign="top" align="center"><p>0.069</p></td>
<td valign="top" align="center"><p>0.307</p></td>
<td valign="top" align="center"><p>12.0</p></td>
</tr>
<tr>
<td valign="top" align="right"><p>Baseline +</p></td>
<td valign="top" align="right"><p>U-Net D.</p></td>
<td valign="top" align="center"><p>0.854</p></td>
<td valign="top" align="center"><p>0.658</p></td>
<td valign="top" align="center"><p>2.950</p></td>
<td valign="top" align="center"><p>0.085</p></td>
<td valign="top" align="center"><p>0.068</p></td>
<td valign="top" align="center"><p>0.299</p></td>
<td valign="top" align="center"><p><bold>28.0</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p>U-Net D. +</p></td>
<td valign="top" align="right"><p><inline-formula><mml:math id="M385" display='block'><mml:msubsup><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mtext>dec</mml:mtext></mml:msub><mml:mtext>cons</mml:mtext></mml:msubsup></mml:math></inline-formula></p></td>
<td valign="top" align="center"><p>0.858</p></td>
<td valign="top" align="center"><p>0.653</p></td>
<td valign="top" align="center"><p>2.860</p></td>
<td valign="top" align="center"><p>0.088</p></td>
<td valign="top" align="center"><p>0.068</p></td>
<td valign="top" align="center"><p>0.296</p></td>
<td valign="top" align="center"><p><bold>19.0</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p><inline-formula><mml:math id="M386" display='block'><mml:msubsup><mml:mi mathvariant="script">L</mml:mi><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mtext>dec</mml:mtext></mml:msub><mml:mtext>cons</mml:mtext></mml:msubsup><mml:mo>+</mml:mo></mml:math></inline-formula></p></td>
<td valign="top" align="right"><p><inline-formula><mml:math id="M387" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>style</mml:mtext></mml:msub></mml:math></inline-formula></p></td>
<td valign="top" align="center"><p>0.831</p></td>
<td valign="top" align="center"><p>0.655</p></td>
<td valign="top" align="center"><p>2.790</p></td>
<td valign="top" align="center"><p>0.091</p></td>
<td valign="top" align="center"><p>0.060</p></td>
<td valign="top" align="center"><p>0.289</p></td>
<td valign="top" align="center"><p><bold>23.0</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p><inline-formula><mml:math id="M388" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:mo>+</mml:mo></mml:math></inline-formula></p></td>
<td valign="top" align="right"><p><inline-formula><mml:math id="M389" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>freq</mml:mtext></mml:msub></mml:math></inline-formula></p></td>
<td valign="top" align="center"><p>0.856</p></td>
<td valign="top" align="center"><p>0.677</p></td>
<td valign="top" align="center"><p>2.410</p></td>
<td valign="top" align="center"><p>0.086</p></td>
<td valign="top" align="center"><p>0.059</p></td>
<td valign="top" align="center"><p>0.288</p></td>
<td valign="top" align="center"><p><bold>20.1</bold></p></td>
</tr>
<tr>
<td valign="top" align="left" rowspan="4"><p><bold>Model</bold></p></td>
<td valign="top" align="right"><p><inline-formula><mml:math id="M390" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>freq</mml:mtext></mml:msub><mml:mo>+</mml:mo></mml:math></inline-formula></p></td>
<td valign="top" align="right"><p>Residual</p></td>
<td valign="top" align="center"><p>0.855</p></td>
<td valign="top" align="center"><p>0.665</p></td>
<td valign="top" align="center"><p>2.310</p></td>
<td valign="top" align="center"><p>0.089</p></td>
<td valign="top" align="center"><p>0.060</p></td>
<td valign="top" align="center"><p>0.285</p></td>
<td valign="top" align="center"><p><bold>25.5</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p>Residual +</p></td>
<td valign="top" align="right"><p>Decoders</p></td>
<td valign="top" align="center"><p>0.863</p></td>
<td valign="top" align="center"><p>0.699</p></td>
<td valign="top" align="center"><p>2.120</p></td>
<td valign="top" align="center"><p>0.079</p></td>
<td valign="top" align="center"><p>0.054</p></td>
<td valign="top" align="center"><p>0.276</p></td>
<td valign="top" align="center"><p><bold>18.5</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p>Decoders +</p></td>
<td valign="top" align="right"><p>Attention</p></td>
<td valign="top" align="center"><p>0.860</p></td>
<td valign="top" align="center"><p>0.665</p></td>
<td valign="top" align="center"><p>2.080</p></td>
<td valign="top" align="center"><p>0.080</p></td>
<td valign="top" align="center"><p>0.059</p></td>
<td valign="top" align="center"><p>0.275</p></td>
<td valign="top" align="center"><p><bold>0.5</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p>Atention +</p></td>
<td valign="top" align="right"><p>ViT</p></td>
<td valign="top" align="center"><p>0.863</p></td>
<td valign="top" align="center"><p>0.665</p></td>
<td valign="top" align="center"><p>2.040</p></td>
<td valign="top" align="center"><p>0.079</p></td>
<td valign="top" align="center"><p>0.057</p></td>
<td valign="top" align="center"><p>0.271</p></td>
<td valign="top" align="center"><p><bold>0.5</bold></p></td>
</tr>
<tr>
<td valign="top" align="left" rowspan="4"><p>Augmentation</p></td>
<td valign="top" align="right"><p>ViT +</p></td>
<td valign="top" align="right"><p>Color</p></td>
<td valign="top" align="center"><p>0.899</p></td>
<td valign="top" align="center"><p>0.692</p></td>
<td valign="top" align="center"><p>1.969</p></td>
<td valign="top" align="center"><p>0.068</p></td>
<td valign="top" align="center"><p>0.054</p></td>
<td valign="top" align="center"><p>0.269</p></td>
<td valign="top" align="center"><p><bold>0.0</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p>Color +</p></td>
<td valign="top" align="right"><p>Rotations</p></td>
<td valign="top" align="center"><p>0.870</p></td>
<td valign="top" align="center"><p>0.682</p></td>
<td valign="top" align="center"><p>2.050</p></td>
<td valign="top" align="center"><p>0.078</p></td>
<td valign="top" align="center"><p>0.055</p></td>
<td valign="top" align="center"><p>0.271</p></td>
<td valign="top" align="center"><p><bold>0.2</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p>Rotations +</p></td>
<td valign="top" align="right"><p>Distortion</p></td>
<td valign="top" align="center"><p>0.876</p></td>
<td valign="top" align="center"><p>0.699</p></td>
<td valign="top" align="center"><p>2.001</p></td>
<td valign="top" align="center"><p>0.074</p></td>
<td valign="top" align="center"><p>0.053</p></td>
<td valign="top" align="center"><p>0.268</p></td>
<td valign="top" align="center"><p><bold>0.0</bold></p></td>
</tr>
<tr>
<td valign="top" align="right"><p>Distortion +</p></td>
<td valign="top" align="right"><p>Erasing</p></td>
<td valign="top" align="center"><p>0.893</p></td>
<td valign="top" align="center"><p>0.727</p></td>
<td valign="top" align="center"><p>1.941</p></td>
<td valign="top" align="center"><p>0.067</p></td>
<td valign="top" align="center"><p>0.052</p></td>
<td valign="top" align="center"><p>0.265</p></td>
<td valign="top" align="center"><p><bold>0.0</bold></p></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="c5-s6.s2">
<label>5.6.2</label>
<title>Qualitative Analysis</title>
<p>We aim to understand which features are exploited by our model for making its predictions. In <xref ref-type="fig" rid="fig-5.7">Figure 5.7</xref>, we show the embeddings of our generator using UMAP. It seems that our model learns to separate between material <italic>families</italic> <target target-type="page" id="pges_133"/>(<italic>e.g. leathers</italic> from <italic>wovens</italic>). Interestingly, the <italic>thickness</italic> of the material is also a relevant parameter. The model is exploiting these semantic patterns without explicit supervision, providing evidence that material microgeometry plays important an important role in its optical appearance.</p>
<fig id="fig-5.7">
<label>Figure 5.7:</label>
<caption><title>Embeddings of the transformer of our generator for data in the training set, reduced using UMAP [<xref ref-type="bibr" rid="CIT266">MHM18</xref>].</title></caption>
<alt-text>Embeddings of the transformer of our generator for data in the training set, reduced using UMAP [MHM18].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.7.jpg"/>
</fig>
</sec>
<sec id="c5-s6.s3">
<label><target target-type="page" id="pges_134"/><target target-type="page" id="pges_135"/>5.6.3</label>
<title>Uncertainty Evaluation</title>
<p>In <xref ref-type="fig" rid="fig-5.8">Figure 5.8</xref>, we show the Pearson correlation matrix between the uncertainty and error for each map, our uncertainty metric (<xref ref-type="disp-formula" rid="Eq005-4">Equation 5.4</xref>), and the render perceptual metric (<xref ref-type="disp-formula" rid="Eq005-5">Equation 5.5</xref>), on our test set. As shown, neither the error nor uncertainties per map explain errors on the render space. Our proposed &#x03C3;<sub>BRDF</sub> achieves a remarkable correlation of 82% with the render error, validating that this metric is useful to predict errors at test time with reasonably high precision.</p>
<fig id="fig-5.8">
<label><target target-type="page" id="pges_136"/>Figure 5.8:</label>
<caption><title>Top-left, correlation between render and pixel-wise losses, and uncertainties. Top-right, plot showing the correlation between our uncertainty metric and the error in render space; the renders illustrate the worst and best cases. Bottom, uncertainty and errors for different material families of the test set.</title></caption>
<alt-text>Top-left, correlation between render and pixel-wise losses, and uncertainties. Top-right, plot showing the correlation between our uncertainty metric and the error in render space; the renders illustrate the worst and best cases. Bottom, uncertainty and errors for different material families of the test set.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.8.jpg"/>
</fig>
<p>In the same plot at the bottom, we distill the uncertainty and render error per family, in which we observe that our model struggles to accurately and confidently predict the reflectance of <italic>satins</italic>, <italic>jacquards</italic>, and <italic>leathers</italic> more than it does for any other structure in our dataset. These structures have complex optical behavior that makes their digitization more challenging (for example, satins exhibit anisotropy, which we do not support in our material model), and are relatively uncommon in our training dataset. <xref ref-type="fig" rid="fig-5.2">Figures 5.2</xref> and <xref ref-type="fig" rid="fig-5.5">5.5</xref> show further examples of our uncertainty estimation for a diverse set of materials.</p>
</sec>
<sec id="c5-s6.s4">
<label>5.6.4</label>
<title>Active Learning</title>
<p>We leverage &#x03C3;<sub>BRDF</sub> for active learning [<xref ref-type="bibr" rid="CIT370">Sol+21</xref>] to identify the samples that contribute most to reduce the error of the model. <xref ref-type="fig" rid="fig-5.9">Figure 5.9</xref> (top) illustrates the process. We start by training a model with 10% of the available training data and measure the uncertainties in the remainder of the dataset. We then select the samples with the highest uncertainties and retrain the model with 20%. We repeat the process with 40, 60, 80 and 100% of the training dataset. For each subset, we compare the performance of this model with four baselines: random-sampling, sampling by the highest uncertainty in normals, roughness, and specular. The results are shown in <xref ref-type="fig" rid="fig-5.9">Figure 5.9</xref> (bottom). Active sampling based on &#x03C3;<sub>BRDF</sub> provides significant gains in sample efficiency, obtaining better accuracy for every map. For instance, an actively trained model with our metric which uses 20% of the training data obtains comparable results to a model trained on three times more (but randomly sampled) data. While &#x03C3;<sub>&#x2221;</sub> is typically more informative than &#x03C3;<sub>rough</sub> and &#x03C3;<sub>spec</sub>, using per-map uncertainties does not provide better results than a random strategy. Finally, <xref ref-type="fig" rid="fig-5.2">Figure 5.2</xref> shows the variation of the probabilistic samples with respect to the ground truth radiance for two materials with high and low uncertainty.</p>
<fig id="fig-5.9">
<label>Figure 5.9:</label>
<caption><title>On top, illustration of our active learning algorithm. On the bottom, results of our active learning experiments. Leveraging &#x03C3;<sub>BRDF</sub> for actively selecting the top-k samples with the highest uncertainty, we achieve better results than a random sampling strategy for every metric we measure.</title></caption>
<alt-text>On top, illustration of our active learning algorithm. On the bottom, results of our active learning experiments. Leveraging &#x03C3;BRDF for actively selecting the top-k samples with the highest uncertainty, we achieve better results than a random sampling strategy for every metric we measure.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.9.jpg"/>
</fig>
</sec>
<sec id="c5-s6.s5">
<label><target target-type="page" id="pges_137"/><target target-type="page" id="pges_138"/>5.6.5</label>
<title>Comparisons with Previous Work</title>
<p>In <xref ref-type="table" rid="c5-tab3">Table 5.3</xref>, we compare our method with previous work on single image material capture. First, we made sure that the training data for these methods included textile materials similar to the ones we choose for testing. Emulating their capture conditions, we took images with a smartphone, with flash and ambient lighting. Note that these capture conditions are not ideal for our method, affecting the final renders if the albedo has shading gradients. However, our goal in this experiment is to evaluate the overall preservation of the material structure in the inferred maps, particularly visible in the normals.</p>
<p>For Shi <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>], we initialize the graph using a fabric material, provided in their open-source implementation and include the metallic map. Our model provides sharper and more accurate estimations without requiring optimization during test. Methods trained on style losses [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>; <xref ref-type="bibr" rid="CIT360">Shi+20</xref>] degrade the semantic structure, while Zhou <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT456">ZK21</xref>] generate similar albedos to ours (note that ours are captured), but provide over-smooth estimations. We provide a comparison of timings and model sizes in <xref ref-type="table" rid="c5-tab2">Table 5.2</xref>. With our efficient model design, we can provide real-time estimations without needing any optimization, which also enables our sampling-based uncertainty quantification.</p>
<table-wrap id="c5-tab2">
<label>Table 5.2:</label>
<caption><title>Model sizes for different methods, and evaluation time in seconds (mean of 100 evaluations on an RTX 2080 GPU), for different output sizes. The methods with <bold>*</bold> use test-time optimization. DiffMat [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>] does not use a pre-trained model.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left"><p><bold>Method</bold></p></th>
<th valign="top" align="left"><p><bold>Size (MB)</bold></p></th>
<th valign="top" align="left"><p><bold>Output Dims</bold></p></th>
<th valign="top" align="left"><p><bold>Eval Time (s)</bold></p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p>Deep Inverse Rendering<bold>*</bold> [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>]</p></td>
<td valign="top" align="left"><p>167</p></td>
<td valign="top" align="left"><p>256x256</p></td>
<td valign="top" align="left"><p>603.5</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Generative Modeling<bold>*</bold> [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>]</p></td>
<td valign="top" align="left"><p>1095.3</p></td>
<td valign="top" align="left"><p>512x384</p></td>
<td valign="top" align="left"><p>218.8</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Diff. Material Graphs<bold>*</bold> [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>]</p></td>
<td valign="top" align="left"><p>-</p></td>
<td valign="top" align="left"><p>512x512</p></td>
<td valign="top" align="left"><p>1209.8</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Adversarial Estimation [<xref ref-type="bibr" rid="CIT456">ZK21</xref>]</p></td>
<td valign="top" align="left"><p>11552.2</p></td>
<td valign="top" align="left"><p>256x256</p></td>
<td valign="top" align="left"><p>0.078</p></td>
</tr>
<tr>
<td valign="middle" align="left"><p><bold>UMat (Ours)</bold></p></td>
<td valign="middle" align="left"><p>22.6</p></td>
<td valign="top" align="left"><p>256x256</p>
<p>512x512</p></td>
<td valign="top" align="left"><p><bold>0.036</bold></p>
<p><bold>0.131</bold></p></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="c5-tab3">
<label><target target-type="page" id="pges_139"/>Table 5.3:</label>
<caption><title>Comparisons of our method with previous work on images captured under different lighting conditions. Top: a smartphone flash-lit image. Middle: a smartphone image with ambient light. Bottom: a flatbed scanner capture. Our method produces the best results preserving the microstructure even when capture conditions degrade due to the sensor resolution. Note that we do not estimate albedos and that absolute intensities for specular and roughness maps are not directly comparable due to differences in the material model.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center"><p>IInput</p></th>
<th valign="top" align="center"><p>Deep Inverse R. [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>]</p></th>
<th valign="top" align="center"><p>Generative Model [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>]</p></th>
<th valign="top" align="center"><p>Diff. Material Graphs [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>]</p></th>
<th valign="top" align="center"><p>Adversarial Est. [<xref ref-type="bibr" rid="CIT456">ZK21</xref>]</p></th>
<th valign="top" align="center"><p><bold>UMat (Ours)</bold></p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center" colspan="6"><p><fig id="fig-6">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.jpg"/>
</fig></p></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="c5-s6.s6">
<label><target target-type="page" id="pges_140"/>5.6.6</label>
<title>Limitations</title>
<p>We show some limitations in <xref ref-type="fig" rid="fig-5.10">Figure 5.10</xref>. The illumination in the scanner hides the wrinkles in the <italic>seersucker</italic> fabric, and our model predicts a flat surface. The <italic>organza</italic> fabric at the bottom is very transparent with visible holes between the yarns. Since the background is white, the model has mistakenly treated the light regions as yarn centers. For the <italic>satin</italic> at the right, the scan image exhibits specular highlights due to the directionality of the yarns. While this image may be problematic to use as an albedo, it does not affect our metrics as we use constant albedos to compute them.</p>
<fig id="fig-5.10">
<label>Figure 5.10:</label>
<caption><title>Limitations cases of our method. On the left, we show a <italic>seersucker</italic>, with wrinkles that are hidden by the diffuse illumination of the device, and a translucent <italic>organza</italic> with holes between the yarns that appear very bright due to the white background of the scanner, and are therefore mistakenly treated as yarn centers. On the right, we show that for highly directional materials, such as <italic>satins</italic>, the diffuse-like illumination in our capture device sometimes introduces specular highlights.</title></caption>
<alt-text>Limitations cases of our method. On the left, we show a seersucker, with wrinkles that are hidden by the diffuse illumination of the device, and a translucent organza with holes between the yarns that appear very bright due to the white background of the scanner, and are therefore mistakenly treated as yarn centers. On the right, we show that for highly directional materials, such as satins, the diffuse-like illumination in our capture device sometimes introduces specular highlights.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.10.jpg"/>
</fig>
</sec>
</sec>
<sec id="c5-s7">
<label>5.7</label>
<title>Conclusions</title>
<p>We have presented a GAN-based method to digitize materials which leverages microgeometry appearance and a flatbed scanner as a capture device. Our method has shown better performance than state-of-the-art solutions that require a single image as input, when it comes to textile materials. To account for potential ambiguities derived from the capture setting, we have presented a method to model the uncertainty in the estimation at test time.</p>
<p>Managing uncertainty in machine learning projects is important to guarantee robust and functional solutions. However, this typically comes at the cost of complex or slow models. In this work, we have presented the first method to quantify uncertainty in single image material digitization, while introducing minimal impact in the training and evaluation processes. While it is currently not possible to discern the source of the uncertainty, whether it is epistemic (uncertainty which can be reduced by increasing the dataset size) or aleatoric (which is derived by a noisy data generation process), our metric has proven useful to identify ambiguous inputs, underrepresented classes, or out-of-distribution data.</p>
<p>We could extend our work in several ways. The most obvious extension is to estimate real albedos, so that we can deal with other types of scanning devices. Further, expanding our material model to give support for more reflectance properties, such as transmittance or anisotropy, could be useful to improve the realism in the render of textiles.</p>
</sec>
<sec id="c5.sA">
<label>5.A</label>
<title>Additional Implementation Details</title>
<sec id="c5.sA.s1">
<label>5.A.1</label>
<title>Model Design</title>
<p>Our model is trained using a GAN framework. In this section, we detail our design choices for the generator and discriminator architecture.</p>
<p><target target-type="page" id="pges_141"/><bold>Generator</bold> For the generator, we use a U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] model, with a few modifications designed to maximize its efficiency, robustness, and generalization capabilities. We specify the full model architecture and layer sizes in <xref ref-type="fig" rid="fig-5.11">Figure 5.11</xref>. We use residual connections [<xref ref-type="bibr" rid="CIT139">He+16</xref>; <xref ref-type="bibr" rid="CIT069">Dia+20</xref>; DLG21] in every convolutional block of the model, for better training convergence and preserving details present in the input images. We use 1 &#x00D7; 1 convolutions on the skip connections. Further, to maximally preserve the appearance and characteristics of every target map, we use a single decoder for each. This has been proposed in different applications, including intrinsic images, material capture, and texture synthesis [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>; <xref ref-type="bibr" rid="CIT339">RG22</xref>; <xref ref-type="bibr" rid="CIT068">DLG21</xref>; <xref ref-type="bibr" rid="CIT456">ZK21</xref>]. To improve the results and enable our uncertainty metric, we append a pixel-wise MLP to each decoder in the model, with Dropout [<xref ref-type="bibr" rid="CIT373">Sri+14</xref>] regularization. Using MLPs after the decoders has been previously explored for material capture [<xref ref-type="bibr" rid="CIT066">Des+19</xref>]. We use Group Normalization [<xref ref-type="bibr" rid="CIT427">WH18</xref>] (with 16 <italic>groups</italic> per layer) and SiLU [<xref ref-type="bibr" rid="CIT086">EUD18</xref>] non-linearities throughout the model. Each convolutional block in the encoder is enhanced with a lightweight Linear Attention module [<xref ref-type="bibr" rid="CIT415">Wan+20b</xref>], with 4 <italic>attention heads</italic>, each with a dimension of 32 hidden units. On the bottleneck, we use a lightweight <italic>MobileViT</italic> Transformer block [<xref ref-type="bibr" rid="CIT268">MR21</xref>], with 128 hidden dimensions for the self-attention and MLPs, 4 layers, and a kernel size of 3. We use a Dropout [<xref ref-type="bibr" rid="CIT097">GG16</xref>] rate of 0.2. We use transposed convolutions for upsampling. Every other implementation detail in the model (strides, bias, poolings) follows [<xref ref-type="bibr" rid="CIT344">RFB15</xref>].</p>
<fig id="fig-5.11">
<label>Figure 5.11:</label>
<caption><title>A full diagram of our generator, including layer sizes and output dimensions for each layer. For Self-Attention, we leverage Linear Attention [<xref ref-type="bibr" rid="CIT415">Wan+20b</xref>], we use a MobileVIT transformer on the bottleneck [<xref ref-type="bibr" rid="CIT268">MR21</xref>], Group Normalization [<xref ref-type="bibr" rid="CIT427">WH18</xref>] and SiLU [<xref ref-type="bibr" rid="CIT086">EUD18</xref>] non-linearities, one decoder per output map and residual connections in every convolutional block. In red, we show the input/output dimensions (spatial, channels) of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, regularizations, and non-linearities.</title></caption>
<alt-text>A full diagram of our generator, including layer sizes and output dimensions for each layer. For Self-Attention, we leverage Linear Attention [Wan+20b], we use a MobileVIT transformer on the bottleneck [MR21], Group Normalization [WH18] and SiLU [EUD18] non-linearities, one decoder per output map and residual connections in every convolu-tional block. In red, we show the input/output dimensions (spatial, channels) of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, regularizations, and non-linearities.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.11.jpg"/>
</fig>
<p><bold>Discriminator</bold> For the discriminator [<xref ref-type="bibr" rid="CIT354">SSK20</xref>], we also use a U-Net [<xref ref-type="bibr" rid="CIT344">RFB15</xref>] model, with a few modifications to improve its performance as a discriminator. We specify the full model architecture and layer sizes in <xref ref-type="fig" rid="fig-5.12">Figure 5.12</xref>. As in the generator, we use residual connections [<xref ref-type="bibr" rid="CIT139">He+16</xref>; <xref ref-type="bibr" rid="CIT069">Dia+20</xref>; DLG21] in every convolutional block of the model, for better training convergence and preserving details present in the input images. We use Spectral Normalization [<xref ref-type="bibr" rid="CIT278">Miy+18</xref>] and SiLU [<xref ref-type="bibr" rid="CIT086">EUD18</xref>] non-linearities throughout the model. On the bottleneck, we use a lightweight <italic>CBAM</italic> attention block [<xref ref-type="bibr" rid="CIT425">Woo+18</xref>]. The single-scalar estimation of the discriminator <inline-formula><mml:math id="M391" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is provided by an MLP with a similar architecture to the <italic>CBAM</italic> module. We initialize the entire model using orthogonal initialization [<xref ref-type="bibr" rid="CIT162">HXP20</xref>]. We use transposed convolutions for upsampling. Every other implementation detail in the model (strides, bias, pooling) follows [<xref ref-type="bibr" rid="CIT344">RFB15</xref>].</p>
<fig id="fig-5.12">
<label>Figure 5.12:</label>
<caption><title>A full diagram of our U-Net residual discriminator, including layer sizes and output dimensions for each layer. We use a CBAM [<xref ref-type="bibr" rid="CIT425">Woo+18</xref>] module on the bottleneck and Spectral Normalization [<xref ref-type="bibr" rid="CIT278">Miy+18</xref>] throughout the network and residual connections in every convolutional block. In red, we show the input/output dimensions of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, non-linearities.</title></caption>
<alt-text>A full diagram of our U-Net residual discriminator, including layer sizes and output dimensions for each layer. We use a CBAM [Woo+18] module on the bottleneck and Spectral Normalization [Miy+18] throughout the network and residual connections in every convolutional block. In red, we show the input/output dimensions of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, non-linearities.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.12.jpg"/>
</fig>
</sec>
<sec id="c5.sA.s2">
<label>5.A.2</label>
<title>Model Training</title>
<p><bold>Optimization</bold> We train the models using PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>] and TorchVision [<xref ref-type="bibr" rid="CIT262">MR10</xref>]. We leverage Kornia for data augmentation [<xref ref-type="bibr" rid="CIT334">Rib+20</xref>]. To accelerate the training process, we leverage mixed precision training and automatic gradient scaling [<xref ref-type="bibr" rid="CIT274">Mic+18</xref>], and train the whole model natively on GPU. Optimization is done using Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>]. Following [<xref ref-type="bibr" rid="CIT354">SSK20</xref>], we use different learning rates for the generator lr = 0.001 and the discriminator lr = 0.005 and a batch size <target target-type="page" id="pges_142"/>of 10, betas= (0.9, 0.99) and a weight decay of 0.000003. We train the models for 100 epochs, which takes 10 hours on an NVIDIA RTX 3060 GPU.</p>
<p><bold>Loss Function</bold> We use the following weights for the loss function: <inline-formula><mml:math id="M392" display='block'><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mspace/><mml:mspace/><mml:mspace/><mml:mspace/><mml:mspace/><mml:mo>)</mml:mo></mml:mrow><mml:mspace/></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>spec</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>rough</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>style</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>0.25</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>freq</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mtext>cons</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn><mml:mo>.</mml:mo></mml:math></inline-formula> For the style loss function, we use the <italic>AlexNet</italic> variant of LPIPS [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>], as it provides a lightweight style loss which has shown success on texture transfer [<xref ref-type="bibr" rid="CIT338">RG21</xref>].</p>
<p><bold>Data Augmentation</bold> Training is done using random crops of 128 &#x00D7; 128 pixels. We perform random rescales uniformly on the 300, 1000 PPI range, with bilinear interpolation. We randomly rotate the SVBRDF on the [&#x2212;10, 10] angle range. We use the algorithm in [<xref ref-type="bibr" rid="CIT338">RG21</xref>] to correctly rotate the normal map. On the HSV color space, we randomly change the properties of the model inputs. We randomly change the hue of the images, selected at random on the whole hue ranges. We also change the saturation, value, and contrast of the input images with factors of [0.9, 1.1], as specified on the Torchvision [<xref ref-type="bibr" rid="CIT262">MR10</xref>] ColorJitter implementation. With a probability of 0.3, we apply random Gaussian noise to the input images <italic>&#x03BC;</italic> = 0, &#x03C3; = 0.10. With a probability of 0.35, we apply random erasing [<xref ref-type="bibr" rid="CIT452">Zho+20</xref>] to the input images. Half of the time, we apply a random Gaussian blur to the input images &#x03C3; &#x223C; <italic>U</italic> (0.1, 5). Finally, with a probability of 0.3, we apply cut-mix data augmentation to the discriminator, as in [<xref ref-type="bibr" rid="CIT354">SSK20</xref>].</p>
<p><bold>Model Evaluation</bold> We use half precision for model evaluation, which makes the process faster and allows us to generate SVBRDF of larger sizes. For images of very high resolutions (eg &#x2265; 4000 &#x00D7; 4000), we use the <italic>stitching</italic> algorithm in [<xref ref-type="bibr" rid="CIT067">DDB20</xref>], with patch sizes of 2048 and stride of 1024 pixels.</p>
</sec>
<sec id="c5.sA.s3">
<label>5.A.3</label>
<title>Artifact Detection</title>
<p>The thresholds for each material map and the kernel size for the uniformity metric have been optimized given a set of 102 manually labeled textures: <inline-formula><mml:math id="M393" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M394" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1.41</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M395" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1.33</mml:mn></mml:math></inline-formula>; <inline-formula><mml:math id="M396" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M397" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.99</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M398" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>3.12</mml:mn></mml:math></inline-formula>. The size of the box filter is <inline-formula><mml:math id="M0399" display='block'><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mtext>Box </mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.1275</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi>DPI</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">X</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. The thresholds for each material map and the kernel size for the uniformity metric have been optimized given a set of 102 manually labeled textures: <inline-formula><mml:math id="M399" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M400" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1.41</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M401" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1.33</mml:mn></mml:math></inline-formula>; <inline-formula><mml:math id="M402" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.01</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M403" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>0.99</mml:mn></mml:math></inline-formula>, <inline-formula><mml:math id="M404" display='block'><mml:msub><mml:mi>t</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>3.12</mml:mn></mml:math></inline-formula>. The size of the box filter is <inline-formula><mml:math id="M405" display='block'><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mtext>Box</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>127.5</mml:mn></mml:math></inline-formula>.</p>
</sec>
<sec id="c5.sA.s4">
<label>5.A.4</label>
<title>Capture Details</title>
<p>We construct the training and testing dataset capturing 10 &#x00D7; 10 cm samples at 1000 PPI using an <italic>EPSON V850 Pro</italic> on its default settings.</p>
<p><target target-type="page" id="pges_143"/>For the comparisons with previous work, we place the fabric samples on a black surface. We use a <italic>Huawei Nova 5T</italic> smartphone, and capture the materials at two distances: a <italic>close-up</italic>, capturing 4 &#x00D7; 4 cms at a distance to the sample of 8 cms, and a <italic>full-size</italic> image, capturing 8 &#x00D7; 8 cms, at a distance to the sample of 13 cms. We use ISO=50, an aperture of <inline-formula><mml:math id="M406" display='block'><mml:mfrac><mml:mi>f</mml:mi><mml:mn>1.8</mml:mn></mml:mfrac></mml:math></inline-formula>, and a focal length of 26mm for every image. For the two distances, we capture the material using ambient illumination, with exposure times of <inline-formula><mml:math id="M407" display='block'><mml:mfrac><mml:mn>1</mml:mn><mml:mn>10</mml:mn></mml:mfrac></mml:math></inline-formula> <italic>s</italic> for the <italic>full size</italic> and of <inline-formula><mml:math id="M408" display='block'><mml:mfrac><mml:mn>1</mml:mn><mml:mn>8</mml:mn></mml:mfrac></mml:math></inline-formula> <italic>s</italic> for the <italic>close-up</italic>. We also capture the images using the smarphone flash lighting, with exposures of <inline-formula><mml:math id="M409" display='block'><mml:mfrac><mml:mn>1</mml:mn><mml:mn>40</mml:mn></mml:mfrac></mml:math></inline-formula> <italic>s</italic> and <inline-formula><mml:math id="M410" display='block'><mml:mfrac><mml:mn>1</mml:mn><mml:mn>50</mml:mn></mml:mfrac></mml:math></inline-formula> <italic>s</italic> for the <italic>full size</italic> and <italic>close-up</italic>, respectively.</p>
</sec>
<sec id="c5.sA.s5">
<label>5.A.5</label>
<title>Comparisons with Previous Work</title>
<p>We perform every comparison with previous work on an RTX NVIDIA 2080, using the default configuration for every method. However, for [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>], we initialize the material graph with a fabric material (fabric suit vintage) provided in their repository, to better match our test data. For [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>], we use the <italic>fine-tuning</italic> configuration.<target target-type="page" id="pges_144"/><target target-type="page" id="pges_145"/></p>
</sec>
</sec>
<sec id="c5.sB">
<label><target target-type="page" id="pges_146"/>5.B</label>
<title>Dataset Analysis</title>
<fig id="fig-5.13">
<label>Figure 5.13:</label>
<caption><title>Visualization of our dataset. We show the percentages of materials in our training dataset, including more detailed subcategories. On their right, we show the average specular and roughness for every category. As shown, there are some structures with distinct characteristics: Satins are highly specular due to the particularities of their yarns, and Piles (eg corduroy) or Plain Weave (eg linen fabrics) are much less glossy. We exploit this relationship between microgeometry and specularity for our estimations.</title></caption>
<alt-text>Visualization of our dataset. We show the percentages of materials in our training dataset, including more detailed subcategories. On their right, we show the average specular and roughness for every category. As shown, there are some structures with distinct characteristics: Satins are highly specular due to the particularities of their yarns, and Piles (eg corduroy) or Plain Weave (eg linen fabrics) are much less glossy. We exploit this relationship between microgeometry and specularity for our estimations.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.13.jpg"/>
</fig>
<fig id="fig-5.14">
<label><target target-type="page" id="pges_147"/>Figure 5.14:</label>
<caption><title>Visualization of some Ground Truth SVBRDF of the different families in our test set. Textiles have very complex and varied microstructures which play an important role on their appearance at different scales.</title></caption>
<alt-text>Visualization of some Ground Truth SVBRDF of the different families in our test set. Textiles have very complex and varied microstructures which play an important role on their appearance at different scales.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.14.jpg"/>
</fig>
</sec>
<sec id="c5.sC">
<label><target target-type="page" id="pges_148"/>5.C</label>
<title>Additional Details And Results</title>
<fig id="fig-5.15">
<label>Figure 5.15:</label>
<caption><title>Additional results of our uncertainty quantification method. On the top, we show a beige <italic>rib</italic> fabric with gray metallic yarns. While there is a low uncertainty for the beige yarns, the metallic yarns are harder to digitize for our model (we do not support metalness in our material model) and it shows a higher uncertainty on those yarns. In the middle, we show a leather material with a very strong structural pattern. Our model shows very low confidence for this material. In the bottom, we show a tartan fabric. Interestingly, our model shows a higher uncertainty on the blue yarns, which are less common than the other yarns in the material.</title></caption>
<alt-text>Additional results of our uncertainty quantification method. On the top, we show a beige rib fabric with gray metallic yarns. While there is a low uncertainty for the beige yarns, the metallic yarns are harder to digitize for our model (we do not support metalness in our material model) and it shows a higher uncertainty on those yarns. In the middle, we show a leather material with a very strong structural pattern. Our model shows very low confidence for this material. In the bottom, we show a tartan fabric. Interestingly, our model shows a higher uncertainty on the blue yarns, which are less common than the other yarns in the material.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.15.jpg"/>
</fig>
<fig id="fig-5.16">
<label><target target-type="page" id="pges_149"/>Figure 5.16:</label>
<caption><title>Additional results of our uncertainty metric. On the top row, we show a plot between our proposed render uncertainty &#x03C3;<sub><italic>BRDF</italic></sub> and the render error, which are highly correlated. We also show average error and uncertainty per family, as well as renders with the lowest and highest errors, compared to the ground truth. Below, we show the same data for normals, specular and roughness errors and uncertainty. There are no correlations between uncertainty and errors for these maps.</title></caption>
<alt-text>Additional results of our uncertainty metric. On the top row, we show a plot between our proposed render uncertainty &#x03C3;BRDF and the render error, which are highly correlated. We also show average error and uncertainty per family, as well as renders with the lowest and highest errors, compared to the ground truth. Below, we show the same data for normals, specular and roughness errors and uncertainty. There are no correlations between uncertainty and errors for these maps.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.16.jpg"/>
</fig>
<fig id="fig-5.17">
<label><target target-type="page" id="pges_150"/>Figure 5.17:</label>
<caption><title>Additional results of our active learning experiment. From top to bottom, we show the renders of ground truth SVBRDFs, the model trained on the 100% of the training data, a model trained on 40% of the training data available, selected following an active learning approach using our uncertainty &#x03C3;<sub><italic>BRDF</italic></sub> as guidance, and a model trained on 40% data, selected randomly. From left to right, we show a very diffuse <italic>Chiffon</italic> fabric, a highly specular <italic>Shantung</italic> fabric and a <italic>Goat Leather</italic> material with very varied microgeometry. In every case, our model trained on 40% of the data following an active learning approach obtains results which are very similar to a model trained on 100% of the data. The model trained on randomly selected 40% data produces highly inaccurate specularity and microgeometry estimations.</title></caption>
<alt-text>Additional results of our active learning experiment. From top to bottom, we show the renders of ground truth SVBRDFs, the model trained on the 100% of the training data, a model trained on 40% of the training data available, selected following an active learning approach using our uncertainty &#x03C3;BRDF as guidance, and a model trained on 40% data, selected randomly. From left to right, we show a very diffuse Chiffon fabric, a highly specular Shantung fabric and a Goat Leather material with very varied microgeometry. In every case, our model trained on 40% of the data following an active learning approach obtains results which are very similar to a model trained on 100% of the data. The model trained on randomly selected 40% data produces highly inaccurate specularity and microgeometry estimations.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.17.jpg"/>
</fig>
<fig id="fig-5.18">
<label><target target-type="page" id="pges_151"/>Figure 5.18:</label>
<caption><title>Results of our method on datasets of previous work. On the top, we show the results of [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>] on their own test set. Our method provides sharper normals which better preserve the structure of the input images. In the middle, we show the results of our method on synthetic rendered data from a material graph from [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>]. Our method provides sharp normals for this synthetic image, which lies outside the distribution of our dataset, composed exclusively of real images. On the bottom, we show an albedo and normals computed using Photometric Stereo [<xref ref-type="bibr" rid="CIT174">Ike81</xref>] for a captured BTF [<xref ref-type="bibr" rid="CIT422">WGK14</xref>], for which our model also provides highly detailed results.</title></caption>
<alt-text>Results of our method on datasets of previous work. On the top, we show the results of [Hen+21] on their own test set. Our method provides sharper normals which better preserve the structure of the input images. In the middle, we show the results of our method on synthetic rendered data from a material graph from [Shi+20]. Our method provides sharp normals for this synthetic image, which lies outside the distribution of our dataset, composed exclusively of real images. On the bottom, we show an albedo and normals computed using Photometric Stereo [Ike81] for a captured BTF [WGK14], for which our model also provides highly detailed results.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-5.18.jpg"/>
</fig>
<table-wrap id="c5-tab4">
<label><target target-type="page" id="pges_152"/>Table 5.4:</label>
<caption><title>Comparisons of our results with previous work on images captured under different conditions for a <italic>curdoroy fabric</italic>: On the first rows, images were captured with a smartphone using the flash image. On the middle rows, using the same smartphone with ambient lighting on different scales. On the final row, a scanner image. Note that for [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>] we use a fabric material for initialization and use their metallic map instead of specular, that we do not estimate albedos and that the material models are not necessarily comparable.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center"><p>Input</p></th>
<th valign="top" align="center"><p>Deep Inverse R. [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>]</p></th>
<th valign="top" align="center"><p>Generative Model [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>]</p></th>
<th valign="top" align="center"><p>Diff. Material Graphs [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>]</p></th>
<th valign="top" align="center"><p>Adversarial Est. [<xref ref-type="bibr" rid="CIT456">ZK21</xref>]</p></th>
<th valign="top" align="center"><p><bold>Our Method</bold></p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center" colspan="6"><p><fig id="fig-7">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.jpg"/>
</fig></p></td>
</tr>
<tr>
<td valign="top" align="center" colspan="6"><p><fig id="fig-8">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-8.jpg"/>
</fig></p></td>
</tr>
<tr>
<td valign="top" align="center" colspan="6"><p><fig id="fig-9">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-9.jpg"/>
</fig></p></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="c5-tab5">
<label><target target-type="page" id="pges_153"/>Table 5.5:</label>
<caption><title>Comparisons of our results with previous work on images captured under different conditions for a <italic>suede leather</italic>: On the first rows, images were captured with a smartphone using the flash image. On the middle rows, using the same smartphone with ambient lighting on different scales. On the final row, a scanner image. Note that for [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>] we use a fabric material for initialization and use their metallic map instead of specular, that we do not estimate albedos and that the material models are not necessarily comparable.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center"><p>Input</p></th>
<th valign="top" align="center"><p>Deep Inverse R. [<xref ref-type="bibr" rid="CIT100">Gao+19</xref>]</p></th>
<th valign="top" align="center"><p>Generative Model [<xref ref-type="bibr" rid="CIT148">Hen+21</xref>]</p></th>
<th valign="top" align="center"><p>Diff. Material Graphs [<xref ref-type="bibr" rid="CIT360">Shi+20</xref>]</p></th>
<th valign="top" align="center"><p>Adversarial Est. [<xref ref-type="bibr" rid="CIT456">ZK21</xref>]</p></th>
<th valign="top" align="center"><p><bold>Our Method</bold></p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center" colspan="6"><p><fig id="fig-10">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-10.jpg"/>
</fig></p></td>
</tr>
<tr>
<td valign="top" align="center" colspan="6"><p><fig id="fig-11">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-11.jpg"/>
</fig></p></td>
</tr>
<tr>
<td valign="top" align="center" colspan="6"><p><fig id="fig-12">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-12.jpg"/>
</fig></p></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</body>
<back>
<fn-group>
<fn id="c5-fn1"><label>1</label> <p>Additional publication details and results are included on <bold>the project website</bold></p></fn>
</fn-group>
</back>
</book-part>
<book-part id="c6" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_154"/><target target-type="page" id="pges_155"/>Chapter 6.</label>
<title>Casual Capture of Fabric Mechanics</title>
</title-group>
</book-part-meta>
<body>
<fig id="fig-6.1">
<label>Figure 6.1:</label>
<caption><title>From two depth images of a fabric sample casually placed in two specific configurations (input), we infer its corresponding set of mechanical parameters. These can be used in a simulator, to visualize the drape of any garment (output). Our method introduces a perceptual drape similarity metric, which enables sorting materials based on their drape.</title></caption>
<alt-text>From two depth images of a fabric sample casually placed in two specific configurations (input), we infer its corresponding set of mechanical parameters. These can be used in a simulator, to visualize the drape of any garment (output). Our method introduces a perceptual drape similarity metric, which enables sorting materials based on their drape.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.1.jpg"/>
</fig>
<p>In this chapter, we propose a method to estimate the mechanical parameters of fabrics using a casual capture setup with a depth camera. Our approach enables to create mechanically-correct digital representations of real-world textile materials, which is a fundamental step for many interactive design and engineering applications. As opposed to existing capture methods, which typically require expensive setups, video sequences, or manual intervention, our solution can capture at scale, is agnostic to the optical appearance of the textile, and facilitates fabric arrangement by non-expert operators. To this end, we propose a sim-to-real strategy to train a learning-based framework that can take as input one or multiple images and outputs a full set of mechanical parameters. Thanks to carefully designed data augmentation and transfer learning protocols, our solution generalizes to real images despite being trained only on synthetic data, hence closing the sim-to-real loop. Key in our work is to demonstrate that evaluating the regression accuracy based on the similarity at parameter space leads to inaccurate distances that do not match the human perception. To overcome this, we propose a novel metric for fabric drape similarity that operates on images instead of the parameter space, allowing us to evaluate our estimation within the context of a similarity rank. We show that our metric correlates with human judgments about the perception of drape similarity, and that our model predictions produce perceptually accurate results compared to the ground truth parameters. The contributions in this chapter led to the following publication<xref ref-type="fn" rid="c6-fn1"><sup>1</sup></xref>:</p>
<fig id="fig-13">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-13.jpg"/>
</fig>
<disp-quote>
<p><target target-type="page" id="pges_156"/>&#x201C;How Will It Drape Like? Capturing Fabric Mechanics From Depth Images&#x201D;</p>
<p>Carlos Rodriguez-Pardo, Melania Prieto-Martin, Dan Casas, Elena Garces</p>
<p><italic>Computer Graphics Forum (Proc. Eurographics 2023)</italic> (2023)</p>
</disp-quote>
<sec id="c6-s1">
<label>6.1</label>
<title>Introduction</title>
<p>Creating accurate digital representations of real-world materials, or <italic>Digital Twins</italic>, is crucial for enabling realistic 3D visualizations suitable for interactive design and predictive engineering. Some industries, like fashion or textile manufacturing, further require these methods to work at a scale to cope with the fast pace of the current production workflows. However, digitizing cloth is challenging due to the high variability and type of fabric samples, where the fabric composition, the microstructure, or the finishing play crucial roles in the perceived appearance.</p>
<p>While casual systems which obtain optical appearance have long been a focus of research, comparably less attention has been paid to estimating <italic>mechanical</italic> properties. Indeed, capturing and simulating the mechanical behavior of cloth is challenging due to the complex interplay between the internal and external forces occurring in this type of physical system, which is highly sensitive to the environmental conditions. Nevertheless, with the current need to create virtual copies instantly, a casual setup able to produce automatic and accurate estimates &#x015B;beyond having a set of <italic>presets</italic> from which to manually choose the closest one&#x015B; of mechanical parameters could prove very valuable.</p>
<p>Previous methods are impractical for scalable and customizable workflows. Accurate fabric parameter acquisition systems require specialized and expensive devices [<xref ref-type="bibr" rid="CIT202">Kaw80</xref>; <xref ref-type="bibr" rid="CIT277">Min95</xref>; <xref ref-type="bibr" rid="CIT052">Cla+90</xref>], which are often slow and need skilled operators. Existing casual capture setups use input video sequences [<xref ref-type="bibr" rid="CIT030">Bha+03</xref>; <xref ref-type="bibr" rid="CIT037">Bou+13</xref>; <xref ref-type="bibr" rid="CIT431">YLL17</xref>] or, even if they take a single image, might require manual user input [<xref ref-type="bibr" rid="CIT188">JC20</xref>]. In this chapter, we present a casual capture system that only requires taking two depth images of the textile posed in a static drape. Our capture setup does not require complex calibration, can be easily manipulated by non-expert operators, and is agnostic to the optical properties of the fabric thanks to leveraging depth images instead of RGB data. <xref ref-type="fig" rid="fig-6.2">Figure 6.2</xref> illustrates <target target-type="page" id="pges_157"/>our capture setup. It involves capturing the fabric with a depth camera in two relaxed positions: the <italic>hanging</italic> scene, whuch conveys the drape when no force other than gravity is applied to it, and the <italic>stretch</italic> scene, which provides cues on the stretching properties of the fabric.</p>
<fig id="fig-6.2">
<label>Figure 6.2:</label>
<caption><title>Capture setup, RGB (top), and depth (bottom) images for <italic>hanging</italic> (left) and <italic>stretch</italic> (right) scenes in rest position. Each scene conveys a different mechanical appearance of the fabric: <italic>hanging</italic> exhibits the overall drape; <italic>stretch</italic> exhibits an extra diagonal tension, which is key to understand the stretching properties.</title></caption>
<alt-text>Capture setup, RGB (top), and depth (bottom) images for hanging (left) and stretch (right) scenes in rest position. Each scene conveys a different mechanical appearance of the fabric: hanging exhibits the overall drape; stretch exhibits an extra diagonal tension, which is key to understand the stretching properties.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.2.jpg"/>
</fig>
<p>We propose a learning-based method to instantly return the mechanical parameters given static depth images and the fabric density as input. Our method relies on a sim-to-real strategy [<xref ref-type="bibr" rid="CIT462">Zhu+21</xref>], leveraging transfer learning and building on a dataset of physical fabrics digitalized with precision equipment. Our model is trained solely with synthetic data, and thanks to carefully designed policies of data augmentation and neural features, it can generalize to real-world scenarios. Under the hood, our approach leverages a custom architecture that enables a flexible design, that can take one or multiple images as input, enhancing its performance when more data is available. We perform an extensive evaluation by means ofablation studies and by measuring aggregated neural network saliency maps, which show that some scenes are more informative than others for predicting the mechanical properties of the fabric. Furthermore, we demonstrate the performance of our model on real-world captured samples, showcasing our system&#x2019;s generalization capabilities.</p>
<p>Key to our work is to demonstrate that evaluating the prediction accuracy of mechanical properties using typical error metrics, such as the Mean Absolute Error (MAE) on the parameter space, leads to inaccurate distances that do not match the human perception. We identify that such mismatches occur due to two factors: first, the parameter space is not bijective &#x015B;i.e., different sets of <target target-type="page" id="pges_158"/>parameters might convey the same drape&#x015B;; and second, a numerical error in a parameter does not necessarily correlate with what we perceive as an error. To address this shortcoming, common in all existing works, and inspired by previous work on similarity metrics for material appearance [<xref ref-type="bibr" rid="CIT225">Lag+19</xref>], natural images [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] or illustration [<xref ref-type="bibr" rid="CIT102">Gar+14</xref>], we propose an image-based metric that measures differences on the mechanical behavior of textiles taking into account the overall drape. We validate that our metric agrees with human perception, and it can be used to sort materials by drape similarity with respect to a reference fabric.</p>
<p>Using our similarity metric, we finally validate that the estimations of our method correlate with human judgments about drape similarity, and that our model predictions produce perceptually accurate results compared to the ground truth parameters. All in all, our approach makes an important step towards solving sim-to-real problems for mechanical estimation since it shows that simulated cloth using inferred parameters maximizes the similarity with respect to real-world target fabrics.</p>
</sec>
<sec id="c6-s2">
<label>6.2</label>
<title>Related Work</title>
<sec id="c6-s2.s1">
<label>6.2.1</label>
<title>Parameter Estimation Methods</title>
<p>Estimating the mechanical properties of real fabric samples is a highly challenging problem for several reasons: the number of uncontrollable extrinsic factors (e.g., wind forces, initial state, collisions, etc.) which affect the predictiveness of the physical simulation; the lack of a standard deformation model and parameter spaces; and the use of computation-intensive simulation methods. Accordingly, a wide range of strategies exist, aimed at overcoming these challenges.</p>
<p><bold>Measurement Devices.</bold> It is common to combine optimization techniques with the output of testing devices to find the optimal set of parameters which best explain the observations [<xref ref-type="bibr" rid="CIT261">Mag+07</xref>; <xref ref-type="bibr" rid="CIT379">SB08</xref>; <xref ref-type="bibr" rid="CIT403">VMF09</xref>; <xref ref-type="bibr" rid="CIT410">WOR11</xref>; <xref ref-type="bibr" rid="CIT275">Mig+12</xref>; <xref ref-type="bibr" rid="CIT054">CTT17</xref>]. Existing technologies of this type are diverse and, as discussed by Kuijpers <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT214">KLG20</xref>], lack of a clear standard. The Kawabata Evaluation System (KES) [<xref ref-type="bibr" rid="CIT202">Kaw80</xref>] is perhaps one of the most well known, measuring 16 coefficients including bending, shearing, and tensile among others. Despite its precision, this method was not widely adopted by the industry due to its lengthy processes and the need for expensive equipment. Consequently, several other methods tried to simplify and unify the methodology with partial success according to some studies [<xref ref-type="bibr" rid="CIT258">LM08</xref>; <xref ref-type="bibr" rid="CIT310">Pow13</xref>]: the Fabric Assurance by Simple Testing (FAST) [<xref ref-type="bibr" rid="CIT277">Min95</xref>], the Fabric Touch Tester (FTT), the CLO Fabric Kit <target target-type="page" id="pges_159"/>2.0, the Fabric Analyser by Browzwear (FAB), the Optitex Mark 10, and the cantilever principle [<xref ref-type="bibr" rid="CIT052">Cla+90</xref>].</p>
<p><bold>Reconstruction-Optimization Methods.</bold> Another set of techniques jointly tackles the reconstruction and parameter optimization problems. By taking as input data from arbitrary real simulations (e.g., the cloth deforming on an avatar [<xref ref-type="bibr" rid="CIT433">Yan+18</xref>]), they iteratively reconstruct and simulate the scene which better explains the observation. Bhat <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT030">Bha+03</xref>] takes as input a video sequence of the cloth and uses simulated annealing to optimize the parameters by measuring its folds. Yang <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT432">YL15</xref>] use multi-view stereo reconstruction to initialize the 3D shape. Runia <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT347">Run+20</xref>] also introduces simulation steps to explain observed phenomena of cloth in the wind. They rely on similarity metrics computed on deep latent spaces to supervise the optimization of the parameters. These methods require simulation steps embedded into the fitting processes, making them computationally expensive due to the high dimensionality of the parameter spaces. Recent differentiable simulation techniques [<xref ref-type="bibr" rid="CIT244">LLK19</xref>; <xref ref-type="bibr" rid="CIT180">JBH19</xref>; <xref ref-type="bibr" rid="CIT166">Hu+19</xref>; <xref ref-type="bibr" rid="CIT287">Mur+20</xref>; <xref ref-type="bibr" rid="CIT237">Li+22</xref>] have proven to be efficient ways to reduce the fitting burden by enabling the computation of gradients with respect to these parameters within latent spaces of neural networks, taking into account dynamics, self-collisions, and contacts.</p>
<p><bold>Data-Driven and Regression Methods.</bold> The third set of methods avoids reconstructing the original 3D scene by working on an estimated feature space and leveraging previously simulated data and machine learning techniques. Our approach falls into this category. Taking videos as input has been explored by Bouman <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT037">Bou+13</xref>] to recover stiffness and area weight using a descriptor of the image based on PCA and optical flow, and later by Yang <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT431">YLL17</xref>], who leverage neural networks to extract image features used for regression. Davis <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT060">Dav+15</xref>] estimated the same simulation parameters by exploiting imperceptible vibrations in high-speed video recordings. Bi <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT031">Bi+18</xref>] further evaluated that humans also need fabric motion to understand its stiffness. Friction coefficients have been estimated using reflectance values [<xref ref-type="bibr" rid="CIT444">ZDN16</xref>] or dynamic videos of cloth sliding through a surface [<xref ref-type="bibr" rid="CIT321">Ras+20</xref>] Instead of regressing the parameters, Huber <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT171">HEW17</xref>] find the most similar cloth in a database using motion descriptors. A different approach only using a single image of the <italic>Cusick drape</italic> was followed by Ju <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT188">JC20</xref>], but it requires a 360<sup>&#x25E6;</sup> scan to reconstruct the target cloth, and a manually fitted Bezier curve to obtain the feature vector. In contrast, we just require a depth map that can be captured easily. Concurrent work [<xref ref-type="bibr" rid="CIT089">Fen+22</xref>] uses multiple-view depth images as input to a trained regressor.</p>
<p>Our approach is inspired by these ideas; however, we do not require optimization &#x015B;providing instant estimation of the parameters&#x015B; and leverage neural features to understand and model fabric behavior in a semi-controlled setup. We demonstrate that our approach works with two images as input to predict <target target-type="page" id="pges_160"/>bending and stretching coefficients without requiring a full video of the piece of fabric.</p>
</sec>
<sec id="c6-s2.s2">
<label>6.2.2</label>
<title>Pre-Trained Models and Transfer Learning</title>
<p>Deep learning models typically require vast amounts of data for generalizing to unseen examples. When this amount of data is not possible to acquire, <italic>transfer learning</italic> techniques help by re-using model parameters trained on a related task [<xref ref-type="bibr" rid="CIT211">KSL19</xref>; <xref ref-type="bibr" rid="CIT314">Rag+19</xref>; <xref ref-type="bibr" rid="CIT035">Bom+21</xref>]. These techniques include: <italic>Fine-tuning</italic> the weights of a pre-trained classification model [<xref ref-type="bibr" rid="CIT423">WKW16</xref>; <xref ref-type="bibr" rid="CIT322">RD17</xref>; <xref ref-type="bibr" rid="CIT049">Che+18</xref>; <xref ref-type="bibr" rid="CIT409">Wan+18</xref>; <xref ref-type="bibr" rid="CIT325">RJ19</xref>; <xref ref-type="bibr" rid="CIT128">Guo+19</xref>; <xref ref-type="bibr" rid="CIT210">Kol+20</xref>]. Pre-training an image descriptor model on contrastive or self-supervised learning tasks, and uses the activations of its last layer as input to the downstream task (<italic>Linear Probing</italic>) [<xref ref-type="bibr" rid="CIT047">Che+20b</xref>; <xref ref-type="bibr" rid="CIT313">Rad+21</xref>; <xref ref-type="bibr" rid="CIT048">CXH21</xref>; <xref ref-type="bibr" rid="CIT138">He+22</xref>; <xref ref-type="bibr" rid="CIT216">Kum+22</xref>]. For domain adaptation problems, it is common to <italic>adapt</italic> the internal representations of pre-trained CNNs so as to efficiency [<xref ref-type="bibr" rid="CIT324">RBV17</xref>; <xref ref-type="bibr" rid="CIT323">RBV18</xref>; <xref ref-type="bibr" rid="CIT342">RB19</xref>; <xref ref-type="bibr" rid="CIT305">Pha+20</xref>; <xref ref-type="bibr" rid="CIT234">LLB22</xref>]. Inspired by these approaches, we design a model that leverages fine-tuning of a pre-trained image CNN classifier as a feature extractor, capable of processing depth images, and extending it to account for additional input variables, and handling multiple images at the same time during inference.</p>
<p><bold>Model Training</bold></p>
<p><bold>Similarity Metric for Drape (<xref ref-type="sec" rid="c6-s1">Section 6</xref>)</bold></p>
<p><bold>Perceptual Evaluation of Fabric Mechanics (<xref ref-type="sec" rid="c7-s1">Section 7</xref>)</bold></p>
</sec>
<sec id="c6-s2.s3">
<label>6.2.3</label>
<title>Similarity Metrics</title>
<p><italic>Full-Reference Image Quality Assessment</italic> (IQA) aims to provide a single score that measures the amount of distortion between two images. Traditionally, these metrics leveraged low-level image statistics. <italic>PSNR</italic> is commonly used for measuring image degradation, but correlates poorly with human perception [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>]. More sophisticated alternatives have been developed, <target target-type="page" id="pges_161"/>including <italic>SSIM</italic>, <italic>mSSIM</italic> [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>], and others [<xref ref-type="bibr" rid="CIT449">Zha+11</xref>; <xref ref-type="bibr" rid="CIT288">Naf+16</xref>; <xref ref-type="bibr" rid="CIT448">ZSL14</xref>; <xref ref-type="bibr" rid="CIT447">ZL12</xref>; <xref ref-type="bibr" rid="CIT328">Rei+18</xref>]. Algorithms based on latent spaces of CNNs [<xref ref-type="bibr" rid="CIT109">GEB16a</xref>] have been extended to better approximate human perception, for example, by training on a large pool of human evaluations <italic>LPIPS</italic> [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>], or by other means [<xref ref-type="bibr" rid="CIT071">Din+20</xref>; <xref ref-type="bibr" rid="CIT311">Pra+18</xref>].</p>
<p>Besides, similarity metrics that measure abstract or complex concepts like style have been proposed for 3D furniture [<xref ref-type="bibr" rid="CIT259">LKS15</xref>], illustration [<xref ref-type="bibr" rid="CIT102">Gar+14</xref>], icons [<xref ref-type="bibr" rid="CIT224">LGG19</xref>], product design [<xref ref-type="bibr" rid="CIT259">LKS15</xref>], or material appearance [<xref ref-type="bibr" rid="CIT225">Lag+19</xref>]. Unlike ours, these metrics require to be trained with human ratings, thus incurring a considerable cost to collect such information via user studies. Instead, our metric does not require specific training, leveraging an off-the-shelf image-based metric. Despite this, we show that our metric correlates with human judgments on the perception of fabric drape similarity, and that can be used to evaluate the overall drape.</p>
</sec>
</sec>
<sec id="c6-s3">
<label>6.3</label>
<title>Overview</title>
<p><xref ref-type="fig" rid="fig-6.3">Figure 6.3</xref> presents an overview of our work. First, in <xref ref-type="sec" rid="c6-s5">Section 6.5</xref>, we introduce our novel solution to infer fabric mechanics directly from depth images. As input, our approach only requires static depth images in two specific configurations, shown in <xref ref-type="fig" rid="fig-6.2">Figure 6.2</xref>, as well as the fabric density which can be easily obtained with conventional equipment. In <xref ref-type="sec" rid="c6-s4">Section 6.4</xref> we describe the datasets of <italic>synthetic</italic> and <italic>real</italic> samples, with mechanical ground truth parameters, used train and evaluate our regressor.</p>
<fig id="fig-6.3">
<label>Figure 6.3:</label>
<caption><title>An overview of the main components of our method. We propose a system to estimate fabric mechanics using depth images of <italic>hanging</italic> and <italic>stretch</italic> scenes as input. To validate the error of our estimations perceptually &#x015B;accounting for the global drape&#x015B;, we propose an image-based drape similarity metric which we validate with human judgments and can be used to sort fabrics by similarity. We show through several metrics that the estimations provided by our method using our similarity metric agree with those given by humans.</title></caption>
<alt-text>An overview of the main components of our method. We propose a system to estimate fabric mechanics using depth images of hanging and stretch scenes as input. To validate the error of our estimations perceptually &#x015B;accounting for the global drape&#x015B;, we propose an image-based drape similarity metric which we validate with human judgments and can be used to sort fabrics by similarity. We show through several metrics that the estimations provided by our method using our similarity metric agree with those given by humans.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.3.jpg"/>
</fig>
<p><target target-type="page" id="pges_162"/>The quantitative evaluation suggests that our method is able to estimate the mechanics within a certain error. However, since a direct interpretation of that error is not human-friendly, we propose a method to evaluate the overall drape in the context of a real scene. In <xref ref-type="sec" rid="c6-s6">Section 6.6</xref>, we introduce our image-based similarity metric for drape, which takes as input renders of the chosen scenes and provides a relative value that is useful to compare the drape of different fabrics. With a user study, explained in <xref ref-type="sec" rid="c6-s7.s1">Section 6.7.1</xref>, we validate that our metric agrees with human preferences on the global perception of fabric drape. Also, in <xref ref-type="sec" rid="c6-s7.s2">Section 6.7.2</xref>, we evaluate our capture method using our drape similarity metric and compare it with human judgments. We effectively validate that our estimations agree with human assessments and provide several qualitative examples in <xref ref-type="sec" rid="c6-s7.s3">Section 6.7.3</xref>.</p>
</sec>
<sec id="c6-s4">
<label>6.4</label>
<title>Datasets</title>
<p>We develop two different datasets of depth images, which we use at different steps of the pipeline to train and evaluate our models: a <italic>synthetic</italic> dataset, generated using physics-based cloth simulation; and <italic>real</italic> dataset, generated using images of real fabric samples. In both datasets, a sample consists of a depth image of a fabric simulated in a scene for which we have the corresponding mechanical parameters, namely: bending [<xref ref-type="bibr" rid="CIT118">Gri+03</xref>] and stretch [<xref ref-type="bibr" rid="CIT403">VMF09</xref>] in the warp, weft, and bias directions and the fabric density, {<italic>kStretchWarp</italic>, <italic>kStretchWeft</italic>, <italic>kStretchBias</italic>, <italic>kBendinghWarp</italic>, <italic>kBendingWeft</italic>, <italic>kBendingBias</italic>, <inline-formula><mml:math id="M411" display='block'><mml:mi>&#x03C1;</mml:mi><mml:mo>}</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mn>7</mml:mn></mml:msup></mml:math></inline-formula>. We support two different static configurations for the scenes: <italic>hanging</italic>, which exhibits the overall drape; and <italic>stretch</italic>, which exhibits diagonal tension and is key to understanding the stretching properties. <xref ref-type="fig" rid="fig-6.2">Figure 6.2</xref> depicts examples of each configuration.</p>
<p><target target-type="page" id="pges_163"/><bold>Simulated Dataset.</bold> To train our model, we generate a synthetic dataset by simulating fabrics in a virtual scenario, replicating the <italic>hanging</italic> and <italic>stretch</italic> configurations. We use a standard simulator, similar to ARCSim [<xref ref-type="bibr" rid="CIT293">NSO12</xref>], with a quadratic strain, a linear strain/stress relationship, and standard definitions for bending and stretch [<xref ref-type="bibr" rid="CIT118">Gri+03</xref>; <xref ref-type="bibr" rid="CIT403">VMF09</xref>]. To model highly anisotropic fabrics, we use three parameters for warp, weft, and bias. Note that the model used for stretch already has some nonlinear behavior (quadratic strain), but more parameters (one per direction) are required to control the nonlinearity of the forces. Thickness is not included as it is implicitly accounted for in the other parameters. Contrarily, density is required as the simulator is dynamic. In a static one, it could be dropped after normalizing the other parameters.</p>
<p>After simulation, we render the resulting mesh (discretized at 5mm/edge) with a white Lambertian material and extract the depth buffer. To ensure that our synthetic dataset covers a wide range of materials, we densely sample the parameter space using a distribution of common mechanical parameters of real-world fabrics. <xref ref-type="fig" rid="fig-6.5">Figure 6.5</xref> shows a sweep of parameters showcasing the variability of resulting drapes in the hanging and stretch scenes. To better understand potential relationships between the parameters, we compute the Spearman correlation <italic>s</italic>, shown in <xref ref-type="fig" rid="fig-6.4">Figure 6.4</xref>, where we observe the higher correlation between the three kStretch coefficients and some correlation between the kBendingBias and the density.</p>
<fig id="fig-6.4">
<label>Figure 6.4:</label>
<caption><title>Spearman correlation matrix between parameters of our synthetic dataset.</title></caption>
<alt-text>Spearman correlation matrix between parameters of our synthetic dataset.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.4.jpg"/>
</fig>
<fig id="fig-6.5">
<label>Figure 6.5:</label>
<caption><title>Sweep of simulation parameters for hanging and stretch scenes. For <italic>kBending</italic>: warp, weft, and bias have the same value, while for <italic>kStretch</italic>, bias changes as 100, 144, 1000.</title></caption>
<alt-text>Sweep of simulation parameters for hanging and stretch scenes. For kBending: warp, weft, and bias have the same value, while for kStretch, bias changes as 100, 144, 1000.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.5.jpg"/>
</fig>
<p><bold>Real Dataset.</bold> To evaluate our model, we test it with real data from images captured by the Intel RealSense SR300. To this end, we casually hang 50 &#x00D7; 50 cm fabric samples using several magnets into a metallic panel, which requires little to no expertise and can be done very fast. Pin location does not need to be centimeter-accurate. Because we use depth images, no special lighting is necessary. We measure fabric area density by weighing a 10 &#x00D7; 10 cm sample and dividing by its area. See <xref ref-type="fig" rid="fig-6.2">Figure 6.2</xref> for a visualization of our capture setup and the accompanying video for an illustration of the process. We capture ten fabrics of diverse compositions and structures, for which we have ground truth mechanical parameters previously obtained with specific equipment and methods [<xref ref-type="bibr" rid="CIT371">Spe+22</xref>].</p>
<sec id="c6-s4.s1">
<label>6.4.1</label>
<title>Data Augmentation</title>
<p>To make our model robust to potential noise and scenario variations common in uncontrolled capture setups (e.g., slightly different camera viewpoint or fabric configuration), we apply several data augmentation strategies to our <italic>synthetic</italic> dataset.</p>
<p><bold>Simulation-Space Data Augmentation.</bold> To enforce robustness to different camera view-points, simulated meshes are rendered within a range of <target target-type="page" id="pges_164"/>different inclinations with respect to the vertical plane. This range covers &#x00B1;5 degrees from the rest position orientation, creating 11 depth maps per material and scene.</p>
<p><bold>Image-Space Data Augmentation.</bold> Real images are largely different from synthetic images, due to distortions, noise, perspective changes, unknown illumination, lens and sensor characteristics, blurs, etc. To enforce the robustness of our models to those distortions, we design an extensive image data augmentation policy consisting of random individual deformations, performed in a particular order. These not only include random noise, blurs, perspective changes and rescales, but also more complex policies such as thin-plate deformations, posterization and erasing. This data augmentation policy bridges the gap between synthetic renders and real depth images, which are typically noisier.</p>
</sec>
</sec>
<sec id="c6-s5">
<label>6.5</label>
<title>Fabric Mechanics from Depth Images</title>
<p>In this section, we present our learning-based approach to estimate fabric mechanical parameters <inline-formula><mml:math id="M412" display='block'><mml:mover><mml:mi mathvariant="script">P</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula><inline-formula><mml:math id="M413" display='block'><mml:mover><mml:mi mathvariant="script">P</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> from depth images. Given a set of depth images, <inline-formula><mml:math id="M414" display='block'><mml:mi mathvariant="script">I</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mtext>hanging</mml:mtext></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mtext>stretch</mml:mtext></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, which depicts the <italic>hanging</italic> and <italic>stretch</italic> scenes, we train a model <inline-formula><mml:math id="M415" display='block'><mml:mi mathvariant="script">M</mml:mi></mml:math></inline-formula> which maps <inline-formula><mml:math id="M416" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula>, along with the</p>
<p>material <italic>density</italic> <italic>&#x03C1;</italic>, to mechanical parameters: <inline-formula><mml:math id="M417" display='block'><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">I</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C1;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mover><mml:mi mathvariant="script">P</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mn>6</mml:mn></mml:msup></mml:math></inline-formula>. To this end, at train time we learn to extract relevant features from depth images, which are then fed into a regressor to learn to predict mechanical features. Importantly, our architecture enables to input sets of images at test time by fusing their respective features. <xref ref-type="fig" rid="fig-6.6">Figure 6.6</xref> illustrates the train and test pipelines.</p>
<fig id="fig-6.6">
<label>Figure 6.6:</label>
<caption><title>Diagram of our training and evaluation pipelines. For <bold>training</bold>, we use a single image of the material, along with its density. The image is processed by our <italic>Feature Extractor</italic>, followed by an MLP, which computes the parameter estimation <inline-formula><mml:math id="M418" display='block'><mml:mover><mml:mi mathvariant="script">P</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> . For <bold>evaluation</bold>, we process each available image with our feature extractor and use a <italic>fusion</italic> operator before feeding it to the trained regressor. We use the same Feature Extractor and MLP for both scenes.</title></caption>
<alt-text>Diagram of our training and evaluation pipelines. For training, we use a single image of the material, along with its density. The image is processed by our Feature Extractor, followed by an MLP, which computes the parameter estimation P^ . For evaluation, we process each available image with our feature extractor and use a fusion operator before feeding it to the trained regressor. We use the same Feature Extractor and MLP for both scenes.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.6.jpg"/>
</fig>
<sec id="c6-s5.s1">
<label><target target-type="page" id="pges_165"/>6.5.1</label>
<title>Neural Network Architecture</title>
<p><bold>Feature Extractor</bold> The first part of the model is a <italic>feature extractor</italic> <inline-formula><mml:math id="M419" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>, which receives as input a single depth image <inline-formula><mml:math id="M420" display='block'><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>I</mml:mi></mml:math></inline-formula> and outputs a feature vector <inline-formula><mml:math id="M421" display='block'><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mfenced open="(" close=")" separators="|"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mrow><mml:mi>sc</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:math></inline-formula> that describes it. This feature extractor is composed of three different components, shown in <xref ref-type="fig" rid="fig-6.6">Figure 6.6</xref>. First, the image is processed by a Convolutional Neural Network (CNN) that outputs a dense latent representation. We use a ResNet-18 [<xref ref-type="bibr" rid="CIT139">He+16</xref>] pre-trained on ImageNet [<xref ref-type="bibr" rid="CIT064">Den+09</xref>], which we finetune with our data. To reduce the gap between real and synthetic images, we equalize the image before feeding it to the feature extractor, during both training and evaluation. Then, the output of this CNN is passed through a <italic>Self-Attention</italic> [<xref ref-type="bibr" rid="CIT443">Zha+19b</xref>] module, which helps the model learn non-local dependencies, enlarging the model receptive field, so it accounts for distant information in the input images. Self-Attention mechanisms were originally designed for language models [<xref ref-type="bibr" rid="CIT398">Vas+17</xref>] but have recently demonstrated significant efficacy for computer vision tasks [<xref ref-type="bibr" rid="CIT318">Ram+19</xref>; <xref ref-type="bibr" rid="CIT443">Zha+19b</xref>; <xref ref-type="bibr" rid="CIT451">ZJK20</xref>; <xref ref-type="bibr" rid="CIT134">Han+22b</xref>]. We add a single Self-Attention layer, as they are expensive to train and evaluate.</p>
<p>Finally, we perform pooling operations to transform the output of the Self-Attention layer to a feature vector, <bold>f</bold><sub>sc</sub>, of a fixed size. In addition to the commonly used max-pooling [<xref ref-type="bibr" rid="CIT364">SZ15</xref>], we further concatenate it with the output of average-pooling, which has been shown to improve the performance of attention modules [<xref ref-type="bibr" rid="CIT453">Zho+16</xref>; <xref ref-type="bibr" rid="CIT161">HSS18</xref>; <xref ref-type="bibr" rid="CIT425">Woo+18</xref>]. The feature vector <bold>f</bold><sub>sc</sub> is thus a concatenation of max-pooled features and average-pooled features: <inline-formula><mml:math id="M422" display='block'><mml:msub><mml:mi mathvariant="bold">f</mml:mi><mml:mi>sc</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi mathvariant="bold">f</mml:mi><mml:mi>sc</mml:mi><mml:mo>max</mml:mo></mml:msubsup><mml:mo>&#x2295;</mml:mo><mml:msubsup><mml:mi mathvariant="bold">f</mml:mi><mml:mi>sc</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msubsup><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p><bold>Fusion</bold> Our design allows us to combine features <bold>f</bold><sub>sc</sub> from more than one scene into a single feature vector <bold>f</bold>. For every image in <italic>I</italic> , we compute <bold>f</bold><sub>sc</sub>. We then <italic>fuse</italic> those feature vectors into a single vector by performing pooling across <bold>f</bold>. Similarly to <bold>f</bold><sub>sc</sub>, <bold>f</bold> is compos types of features: <inline-formula><mml:math id="M423" display='block'><mml:mi>f</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:munder><mml:mo movablelimits="true">max</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true">max</mml:mo></mml:mrow></mml:msubsup><mml:mo>}</mml:mo></mml:mrow><mml:mo>&#x2295;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msubsup><mml:mo>}</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> The max-pooled features are fused using the maximum value across scenes, while the average-pooled features are fused using the mean across scenes. We illustrate this procedure in <xref ref-type="fig" rid="fig-6.6">Figure 6.6</xref>.</p>
<p><bold>Parameter Regressor</bold> Our last component is a fully-connected Multi-Layer Perceptron (MLP), which takes as input the feature vector <bold>f</bold> and the material density, <italic>&#x03C1;</italic>, and outputs the simulation parameters <inline-formula><mml:math id="M426" display='block'><mml:mover><mml:mi mathvariant="script">P</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> . As our loss function, we compare the real parameters <inline-formula><mml:math id="M427" display='block'><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula> with the model estimations <inline-formula><mml:math id="M428" display='block'><mml:mi mathvariant="script">M</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="script">I</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C1;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mover><mml:mi mathvariant="script">P</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> using an <italic>&#x2113;</italic><sub>2</sub> norm.</p>
</sec>
<sec id="c6-s5.s2">
<label>6.5.2</label>
<title>Quantitative Evaluation</title>
<p>In this section, we quantitatively evaluate the performance of the method for estimating the mechanical parameters. We validate the design choices of the model and evaluate our results depending on the type of input used.<target target-type="page" id="pges_166"/></p>
<sec id="c6-s5.s2.1">
<title><target target-type="page" id="pges_167"/>Ablation Study of the Model Design</title>
<p>We aim to understand the effective contribution of the data augmentation strategy, and the network architecture design. For these experiments, we randomly split the synthetic data in 90% for training and 10% for validation, using the same split for every experiment. Results are shown in <xref ref-type="table" rid="c6-tab1">Table 6.1</xref>. As our baseline, we use a simple model without <italic>self-attention</italic>, without <italic>average-pooling</italic> features, where the CNN backbone is randomly initialized, and without data augmentation. From this baseline, we progressively add different components and measure the performance of the validation data using Mean Absolute Error (MAE) and Spearman correlation (<italic>s</italic>). The parameters are normalized using the minimum and maximum values of the training set.</p>
<table-wrap id="c6-tab1">
<label>Table 6.1:</label>
<caption><title>Ablation study of the neural architecture and data augmentation. From left to right, we build upon our baseline and progressively add: simulation-space data augmentation, image-space data augmentation, pre-training, self-attention, and average-pooling. On both MAE (<italic>&#x2113;</italic><sub>1</sub>) and correlation (<italic>s</italic>) metrics, we observe increased performance on the validation set in every added component. Using a pre-trained network for feature extraction yields the largest gains. We use a color code to highlight <bold>best</bold> and <bold>worst</bold> cases.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center" rowspan="2"><p><bold>Metric</bold></p></th>
<th valign="top" align="center" rowspan="2"><p><bold>Parameter</bold></p></th>
<th valign="top" align="center" rowspan="2"><p><bold>Baseline</bold></p></th>
<th valign="top" align="center" colspan="2"><p><bold>Data Augmentation</bold></p></th>
<th valign="top" align="center" colspan="3"><p><bold>Architecture</bold></p></th>
</tr>
<tr>
<th valign="top" align="center"><p>w/ sim</p></th>
<th valign="top" align="center"><p>w/ image</p></th>
<th valign="top" align="center"><p>w/ pre-Train</p></th>
<th valign="top" align="center"><p>w/ attention</p></th>
<th valign="top" align="center"><p>w/ pooling</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="right" rowspan="8"><p><inline-formula><mml:math id="M424" display='block'><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></p></td>
<td valign="top" align="center"><p>kStretchWeft</p></td>
<td valign="top" align="center"><p>0.122</p></td>
<td valign="top" align="center"><p>0.118</p></td>
<td valign="top" align="center"><p>0.112</p></td>
<td valign="top" align="center"><p>0.087</p></td>
<td valign="top" align="center"><p>0.071</p></td>
<td valign="top" align="center"><p>0.071</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kStretchWarp</p></td>
<td valign="top" align="center"><p>0.109</p></td>
<td valign="top" align="center"><p>0.101</p></td>
<td valign="top" align="center"><p>0.102</p></td>
<td valign="top" align="center"><p>0.074</p></td>
<td valign="top" align="center"><p>0.066</p></td>
<td valign="top" align="center"><p>0.061</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kStretchBias</p></td>
<td valign="top" align="center"><p>0.050</p></td>
<td valign="top" align="center"><p>0.043</p></td>
<td valign="top" align="center"><p>0.046</p></td>
<td valign="top" align="center"><p>0.043</p></td>
<td valign="top" align="center"><p>0.042</p></td>
<td valign="top" align="center"><p>0.038</p></td>
</tr>
<tr>
<td valign="top" align="center"><p><bold>Avg. Stretch</bold></p></td>
<td valign="top" align="center"><p>0.094</p></td>
<td valign="top" align="center"><p>0.087</p></td>
<td valign="top" align="center"><p>0.087</p></td>
<td valign="top" align="center"><p>0.068</p></td>
<td valign="top" align="center"><p>0.060</p></td>
<td valign="top" align="center"><p>0.057</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingWeft</p></td>
<td valign="top" align="center"><p>0.095</p></td>
<td valign="top" align="center"><p>0.093</p></td>
<td valign="top" align="center"><p>0.091</p></td>
<td valign="top" align="center"><p>0.084</p></td>
<td valign="top" align="center"><p>0.0.82</p></td>
<td valign="top" align="center"><p>0.072</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingWarp</p></td>
<td valign="top" align="center"><p>0.124</p></td>
<td valign="top" align="center"><p>0.119</p></td>
<td valign="top" align="center"><p>0.112</p></td>
<td valign="top" align="center"><p>0.110</p></td>
<td valign="top" align="center"><p>0.101</p></td>
<td valign="top" align="center"><p>0.094</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingBias</p></td>
<td valign="top" align="center"><p>0.086</p></td>
<td valign="top" align="center"><p>0.081</p></td>
<td valign="top" align="center"><p>0.086</p></td>
<td valign="top" align="center"><p>0.081</p></td>
<td valign="top" align="center"><p>0.068</p></td>
<td valign="top" align="center"><p>0.060</p></td>
</tr>
<tr>
<td valign="top" align="center"><p><bold>Avg. Bending</bold></p></td>
<td valign="top" align="center"><p>0.102</p></td>
<td valign="top" align="center"><p>0.098</p></td>
<td valign="top" align="center"><p>0.097</p></td>
<td valign="top" align="center"><p>0.092</p></td>
<td valign="top" align="center"><p>0.084</p></td>
<td valign="top" align="center"><p>0.075</p></td>
</tr>
<tr>
<td valign="top" align="right" rowspan="8"><p><inline-formula><mml:math id="M425" display='block'><mml:mi>r</mml:mi><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></p></td>
<td valign="top" align="center"><p>kStretchWeft</p></td>
<td valign="top" align="center"><p>0.445</p></td>
<td valign="top" align="center"><p>0.475</p></td>
<td valign="top" align="center"><p>0.453</p></td>
<td valign="top" align="center"><p>0.643</p></td>
<td valign="top" align="center"><p>0.788</p></td>
<td valign="top" align="center"><p>0.798</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kStretchWarp</p></td>
<td valign="top" align="center"><p>0.543</p></td>
<td valign="top" align="center"><p>0.503</p></td>
<td valign="top" align="center"><p>0.510</p></td>
<td valign="top" align="center"><p>0.645</p></td>
<td valign="top" align="center"><p>0.715</p></td>
<td valign="top" align="center"><p>0.771</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kStretchBias</p></td>
<td valign="top" align="center"><p>0.477</p></td>
<td valign="top" align="center"><p>0.508</p></td>
<td valign="top" align="center"><p>0.608</p></td>
<td valign="top" align="center"><p>0.701</p></td>
<td valign="top" align="center"><p>0.733</p></td>
<td valign="top" align="center"><p>0.781</p></td>
</tr>
<tr>
<td valign="top" align="center"><p><bold>Avg. Stretch</bold></p></td>
<td valign="top" align="center"><p>0.488</p></td>
<td valign="top" align="center"><p>0.495</p></td>
<td valign="top" align="center"><p>0.523</p></td>
<td valign="top" align="center"><p>0.660</p></td>
<td valign="top" align="center"><p>0.745</p></td>
<td valign="top" align="center"><p>0.783</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingWeft</p></td>
<td valign="top" align="center"><p>0.611</p></td>
<td valign="top" align="center"><p>0.684</p></td>
<td valign="top" align="center"><p>0.781</p></td>
<td valign="top" align="center"><p>0.783</p></td>
<td valign="top" align="center"><p>0.798</p></td>
<td valign="top" align="center"><p>0.863</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingWarp</p></td>
<td valign="top" align="center"><p>0.614</p></td>
<td valign="top" align="center"><p>0.654</p></td>
<td valign="top" align="center"><p>0.747</p></td>
<td valign="top" align="center"><p>0.772</p></td>
<td valign="top" align="center"><p>0.806</p></td>
<td valign="top" align="center"><p>0.921</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingBias</p></td>
<td valign="top" align="center"><p>0.624</p></td>
<td valign="top" align="center"><p>0.828</p></td>
<td valign="top" align="center"><p>0.837</p></td>
<td valign="top" align="center"><p>0.872</p></td>
<td valign="top" align="center"><p>0.893</p></td>
<td valign="top" align="center"><p>0.942</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Avg. Bending</p></td>
<td valign="top" align="center"><p>0.616</p></td>
<td valign="top" align="center"><p>0.722</p></td>
<td valign="top" align="center"><p>0.788</p></td>
<td valign="top" align="center"><p>0.809</p></td>
<td valign="top" align="center"><p>0.832</p></td>
<td valign="top" align="center"><p>0.909</p></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Given the training configuration with all the data augmentation &#x015B;which provides a small increase in performance most likely because the validation dataset is synthetic data&#x015B; we evaluate the neural architecture.</p>
<p>Using a CNN backbone pre-trained on ImageNet [<xref ref-type="bibr" rid="CIT064">Den+09</xref>] instead of a randomly initialized one, we observe a significant increase in model performance across every parameter and metric. Training a feature extractor that receives images from both scenes at the same time would not allow us to leverage pre-training, which would negatively impact generalization. Then, we add a <italic>self-attention</italic> [<xref ref-type="bibr" rid="CIT443">Zha+19b</xref>] layer after the CNN backbone, which allows the model to integrate information that is present in distant areas of the images. Interestingly, this module significantly helps predict the <italic>kStretch</italic> parameters while having a more minor influence on the <italic>kBending</italic>. Finally, adding <italic>average-pooling</italic> in addition to the commonly used <italic>max-pooling</italic> has a highly positive impact on the error rates. We use this last configuration with all the components for all the results shown in the chapter. It is worth noting that the <italic>kBendingBias</italic> is easily predicted by the model, even in its most basic configuration. This is likely because this parameter correlates most strongly with the density of the material (shown in <xref ref-type="fig" rid="fig-6.4">Figure 6.4</xref>), so the model can leverage this information for the predictions.</p>
</sec>
<sec id="c6-s5.s2.2">
<title>Evaluation of Input Influence</title>
<p>The design of our method supports taking as input one or multiple images. In this experiment, presented in <xref ref-type="table" rid="c6-tab2">Table 6.2</xref>, we evaluate the error testing different configurations of the input. Note that we train a different model for each configuration. In every case, using the <italic>density</italic> as input helps the model to generalize. This is particularly relevant for <italic>kBending</italic> parameters, for which the <italic>density</italic> alone provides more information than depth images. The <italic>stretch</italic> scene is typically more informative than the <italic>bending</italic> one, as both MAE and correlations are usually better when it is provided. When using both scenes simultaneously, the model provides moreaccurate estimations than any of the scenes individually, showing that the two scenes provide complementary information.</p>
<table-wrap id="c6-tab2">
<label><target target-type="page" id="pges_168"/>Table 6.2:</label>
<caption><title>Results for real depth images varying the input. From left to right, the input is: only density, only depth images, and both density and depth. The best results are obtained using every data source as input. For <italic>kBending</italic>, the <italic>density</italic> alone provides more information than only depth images. We use a color code to highlight <bold>best</bold> and <bold>worst</bold> cases.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center" rowspan="2"><p><bold>Metric</bold></p></th>
<th valign="top" align="center" rowspan="2"><p><bold>Parameter</bold></p></th>
<th valign="top" align="center" rowspan="2"><p><bold>Density</bold></p></th>
<th valign="top" align="center" colspan="3"><p><bold>OnlyDepth</bold></p></th>
<th valign="top" align="center" colspan="3"><p><bold>Density &#x0026; Depth</bold></p></th>
</tr>
<tr>
<th valign="top" align="center"><bold>Stretch</bold></th>
<th valign="top" align="center"><bold>Hanging</bold></th>
<th valign="top" align="center"><bold>Both</bold></th>
<th valign="top" align="center"><bold>Stretch</bold></th>
<th valign="top" align="center"><bold>Hanging</bold></th>
<th valign="top" align="center"><bold>Both</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="right" rowspan="8"><inline-formula><mml:math id="M429" display='block'><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></td>
<td valign="top" align="center"><p>kStretchWeft</p></td>
<td valign="top" align="center"><p>0.113</p></td>
<td valign="top" align="center"><p>0.102</p></td>
<td valign="top" align="center"><p>0.107</p></td>
<td valign="top" align="center"><p>0.062</p></td>
<td valign="top" align="center"><p>0.054</p></td>
<td valign="top" align="center"><p>0.056</p></td>
<td valign="top" align="center"><p>0.051</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kStretchWarp</p></td>
<td valign="top" align="center"><p>0.091</p></td>
<td valign="top" align="center"><p>0.067</p></td>
<td valign="top" align="center"><p>0.081</p></td>
<td valign="top" align="center"><p>0.056</p></td>
<td valign="top" align="center"><p>0.055</p></td>
<td valign="top" align="center"><p>0.059</p></td>
<td valign="top" align="center"><p>0.052</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kStretchBias</p></td>
<td valign="top" align="center"><p>0.034</p></td>
<td valign="top" align="center"><p>0.039</p></td>
<td valign="top" align="center"><p>0.054</p></td>
<td valign="top" align="center"><p>0.034</p></td>
<td valign="top" align="center"><p>0.033</p></td>
<td valign="top" align="center"><p>0.036</p></td>
<td valign="top" align="center"><p>0.031</p></td>
</tr>
<tr>
<td valign="top" align="center"><p><bold>Mean Stretch</bold></p></td>
<td valign="top" align="center"><p>0.079</p></td>
<td valign="top" align="center"><p>0.069</p></td>
<td valign="top" align="center"><p>0.081</p></td>
<td valign="top" align="center"><p>0.051</p></td>
<td valign="top" align="center"><p>0.047</p></td>
<td valign="top" align="center"><p>0.050</p></td>
<td valign="top" align="center"><p>0.045</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingWeft</p></td>
<td valign="top" align="center"><p>0.142</p></td>
<td valign="top" align="center"><p>0.233</p></td>
<td valign="top" align="center"><p>0.213</p></td>
<td valign="top" align="center"><p>0.145</p></td>
<td valign="top" align="center"><p>0.128</p></td>
<td valign="top" align="center"><p>0.139</p></td>
<td valign="top" align="center"><p>0.125</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingWarp</p></td>
<td valign="top" align="center"><p>0.126</p></td>
<td valign="top" align="center"><p>0.184</p></td>
<td valign="top" align="center"><p>0.126</p></td>
<td valign="top" align="center"><p>0.074</p></td>
<td valign="top" align="center"><p>0.037</p></td>
<td valign="top" align="center"><p>0.063</p></td>
<td valign="top" align="center"><p>0.035</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingBias</p></td>
<td valign="top" align="center"><p>0.094</p></td>
<td valign="top" align="center"><p>0.166</p></td>
<td valign="top" align="center"><p>0.081</p></td>
<td valign="top" align="center"><p>0.054</p></td>
<td valign="top" align="center"><p>0.055</p></td>
<td valign="top" align="center"><p>0.069</p></td>
<td valign="top" align="center"><p>0.046</p></td>
</tr>
<tr>
<td valign="top" align="center"><p><bold>Mean Bending</bold></p></td>
<td valign="top" align="center"><p>0.121</p></td>
<td valign="top" align="center"><p>0.194</p></td>
<td valign="top" align="center"><p>0.140</p></td>
<td valign="top" align="center"><p>0.091</p></td>
<td valign="top" align="center"><p>0.073</p></td>
<td valign="top" align="center"><p>0.090</p></td>
<td valign="top" align="center"><p>0.069</p></td>
</tr>
<tr>
<td valign="top" align="right" rowspan="8"><p><inline-formula><mml:math id="M430" display='block'><mml:mi>r</mml:mi><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></p></td>
<td valign="top" align="center"><p>kStretchWeft</p></td>
<td valign="top" align="center"><p>0.184</p></td>
<td valign="top" align="center"><p>0.407</p></td>
<td valign="top" align="center"><p>0.403</p></td>
<td valign="top" align="center"><p>0.418</p></td>
<td valign="top" align="center"><p>0.712</p></td>
<td valign="top" align="center"><p>0.469</p></td>
<td valign="top" align="center"><p>0.728</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kStretchWarp</p></td>
<td valign="top" align="center"><p>0.002</p></td>
<td valign="top" align="center"><p>0.502</p></td>
<td valign="top" align="center"><p>0.407</p></td>
<td valign="top" align="center"><p>0.503</p></td>
<td valign="top" align="center"><p>0.520</p></td>
<td valign="top" align="center"><p>0.433</p></td>
<td valign="top" align="center"><p>0.533</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kStretchBias</p></td>
<td valign="top" align="center"><p>0.289</p></td>
<td valign="top" align="center"><p>-0.141</p></td>
<td valign="top" align="center"><p>-0.06
8</p></td>
<td valign="top" align="center"><p>-0.00
7</p></td>
<td valign="top" align="center"><p>0.367</p></td>
<td valign="top" align="center"><p>0.383</p></td>
<td valign="top" align="center"><p>0.550</p></td>
</tr>
<tr>
<td valign="top" align="center"><p><bold>Mean Stretch</bold></p></td>
<td valign="top" align="center"><p>0.158</p></td>
<td valign="top" align="center"><p>0.256</p></td>
<td valign="top" align="center"><p>0.247</p></td>
<td valign="top" align="center"><p>0.305</p></td>
<td valign="top" align="center"><p>0.533</p></td>
<td valign="top" align="center"><p>0.428</p></td>
<td valign="top" align="center"><p>0.604</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingWeft</p></td>
<td valign="top" align="center"><p>0.483</p></td>
<td valign="top" align="center"><p>0.052</p></td>
<td valign="top" align="center"><p>0.267</p></td>
<td valign="top" align="center"><p>0.625</p></td>
<td valign="top" align="center"><p>0.673</p></td>
<td valign="top" align="center"><p>0.683</p></td>
<td valign="top" align="center"><p>0.717</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingWarp</p></td>
<td valign="top" align="center"><p>0.317</p></td>
<td valign="top" align="center"><p>0.048</p></td>
<td valign="top" align="center"><p>0.156</p></td>
<td valign="top" align="center"><p>0.407</p></td>
<td valign="top" align="center"><p>0.467</p></td>
<td valign="top" align="center"><p>0.433</p></td>
<td valign="top" align="center"><p>0.533</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>kBendingBias</p></td>
<td valign="top" align="center"><p>0.357</p></td>
<td valign="top" align="center"><p>-0.044</p></td>
<td valign="top" align="center"><p>0.108</p></td>
<td valign="top" align="center"><p>0.250</p></td>
<td valign="top" align="center"><p>0.333</p></td>
<td valign="top" align="center"><p>0.417</p></td>
<td valign="top" align="center"><p>0.546</p></td>
</tr>
<tr>
<td valign="top" align="center"><p><bold>Mean Bending</bold></p></td>
<td valign="top" align="center"><p>0.386</p></td>
<td valign="top" align="center"><p>0.019</p></td>
<td valign="top" align="center"><p>0.177</p></td>
<td valign="top" align="center"><p>0.427</p></td>
<td valign="top" align="center"><p>0.491</p></td>
<td valign="top" align="center"><p>0.511</p></td>
<td valign="top" align="center"><p>0.599</p></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="c6-s5.s2.3">
<title><target target-type="page" id="pges_169"/>Neural Saliency Maps</title>
<p>We aim to understand which parts of the scenes are most relevant for the model when making its predictions. To do so, we create saliency maps using <italic>Axion-Based Class-Activation Mappings</italic> [<xref ref-type="bibr" rid="CIT096">Fu+20</xref>], which averages the activations of the deep layers weighted by their importance with respect to each target parameter, and aggregate them over our real dataset. In <xref ref-type="fig" rid="fig-6.7">Figure 6.7</xref> we observe that each scene provides the model with different cues. For <italic>kStretch</italic> parameters, the model is sensitive to the central wrinkle of the <italic>stretch</italic> scene and the central fold of the <italic>hanging</italic>. For <italic>kBending</italic> parameters, the model is sensitive to areas in the borders of the fabric where small but noticeable wrinkles are present, in both scenes. As in other experiments, we observe that the model can find more relevant features on the <italic>stretch</italic> scene.</p>
<fig id="fig-6.7">
<label>Figure 6.7:</label>
<caption><title>Saliency maps [<xref ref-type="bibr" rid="CIT096">Fu+20</xref>] aggregated per parameter. The model relies on the central areas of the fabric samples for predicting the stretch parameters. For bending, it is most sensitive to areas on the borders of the samples.</title></caption>
<alt-text>Saliency maps [Fu+20] aggregated per parameter. The model relies on the central areas of the fabric samples for predicting the stretch parameters. For bending, it is most sensitive to areas on the borders of the samples.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.7.jpg"/>
</fig>
</sec>
</sec>
</sec>
<sec id="c6-s6">
<label>6.6</label>
<title>A Similarity Metric for Drape</title>
<p>The regressor introduced in <xref ref-type="sec" rid="c6-s5">Section 6.5</xref> allows us to infer the mechanical parameters of a target fabric. However, since a direct interpretation of such parametric space is not human-friendly, it is very difficult to understand the residual errors shown in <xref ref-type="table" rid="c6-tab1">Tables 6.1</xref> and <xref ref-type="table" rid="c6-tab2">6.2</xref>. Do the regressed parameters produce a drape similar to the target image? Notice that, since the mechanical parameter space is non-orthogonal, small parameter changes may produce unexpected deformations. Therefore, we hypothesize that a <italic>perceptual</italic> similarity metric for drape is needed to interpret our quantitative results. We describe this metric next.</p>
<sec id="c6-s6.s1">
<label>6.6.1</label>
<title>Image-based Similarity of Drape</title>
<p>Motivated by our hypothesis that the <italic>hanging</italic> and <italic>stretch</italic> scenes are sufficient to convey the fabric mechanics, we propose an image-based similarity <target target-type="page" id="pges_170"/>metric using renders of such scenes. Let <inline-formula><mml:math id="M431" display='block'><mml:mi mathvariant="script">P</mml:mi><mml:mo>&#x2286;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mn>7</mml:mn></mml:msup></mml:math></inline-formula> be the parameter space of our simulator, <inline-formula><mml:math id="M432" display='block'><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M433" display='block'><mml:msub><mml:mi>P</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula> two different parameter sets, and <italic>a</italic> &#x223C; <inline-formula><mml:math id="M434" display='block'><mml:mi mathvariant="script">R</mml:mi></mml:math></inline-formula>(<italic>P<sub>a</sub></italic>, scene) and <italic>b</italic> &#x223C; <inline-formula><mml:math id="M435" display='block'><mml:mi mathvariant="script">R</mml:mi></mml:math></inline-formula>(<italic>P<sub>b</sub></italic>, scene) two rendered simulations obtained for a certain scene configuration. We define a distance metric for a particular scene as:</p>
<p><disp-formula id="Eq006-1"><label>(6.1)</label> <mml:math id="M436" display='block'><mml:msub><mml:mi>d</mml:mi><mml:mtext>scene</mml:mtext></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mi>IM</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mfrac></mml:math></disp-formula></p>
<p>where IM is an image-space distance metric, and <italic>N</italic> is the number of different simulations we run. Since real cloth is very sensitive to parameters such as initial state or initial shape, in order to learn a metric that is robust to real-world conditions, we perturb the initial state and boundary conditions in a set of simulations using random jittering to the initial forces. We</p>
<p>empirically found that averaging over multiple simulations for the same set of (<italic>P<sub>a</sub></italic>, <italic>P<sub>b</sub></italic>) gives us a more informative metric.</p>
<p>Furthermore, we take into account both <italic>hanging</italic> and <italic>stretch</italic> scenes, hence our final metric is defined by averaging their distances across both scenarios, resulting in our final metric:</p>
<p><disp-formula id="Eq006-2"><label>(6.2)</label> <mml:math id="M437" display='block'><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mtext>hanging</mml:mtext></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mtext>stretch</mml:mtext></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:math></disp-formula></p>
<p>Note that we propose a similarity metric, which, as opposed to real distance metrics, does not necessarily have to meet the metric axioms [<xref ref-type="bibr" rid="CIT393">TK74</xref>]: it can produce asymmetric values, violate the triangle inequality, and does not need to define what the identity is. According to <xref ref-type="disp-formula" rid="Eq006-1">Equation 6.1</xref>, the distance of one fabric with itself is not necessarily zero; it just needs to satisfy a minimum requirement where the distance of every material with itself should be smaller than the distance of any material with any other material,</p>
<p><disp-formula id="Eq006-3"><label>(6.3)</label> <mml:math id="M438" display='block'><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x2260;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></disp-formula></p>
<p>In order to remove the possible influence of optical properties, scene illumination, and camera parameters, we use the same scene configuration for every render, with grayscale albedo and a Lambertian BRDF.</p>
<p>For IM we use LPIPS [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] which we empirically found to perform better than other alternative image metrics. We find that metrics based on pre-trained neural networks work better than lower-level alternatives, while content-aware distances are more powerful for this purpose than style-aware metrics. This suggests that the size, position, and shape of the wrinkles and deformations of the fabrics are important factors that explain differences between materials. We empirically found that <italic>N</italic> = 5 simulations are typically enough, as more samples provide very marginal improvements. See the supplementary material of this chapter for more details about the proposed metric.</p>
</sec>
</sec>
<sec id="c6-s7">
<label><target target-type="page" id="pges_171"/>6.7</label>
<title>Evaluation</title>
<p>We propose to evaluate our method from <xref ref-type="sec" rid="c6-s5">Section 6.5</xref> and the metric from <xref ref-type="sec" rid="c6-s6">Section 6.6</xref> by comparing our estimations with human ratings (recall that we provide quantitative errors per parameter in <xref ref-type="sec" rid="c6-s5.s2">Section 6.5.2</xref>). To this end, in <xref ref-type="sec" rid="c6-s7.s1">Section 6.7.1</xref>, we first collect a large number of ground truth human judgments about the similarity of triplets of fabrics. Then, in <xref ref-type="sec" rid="c6-s7.s2">Section 6.7.2</xref>, we demonstrate that our image-based metric using ground truth parameters, as well as the estimated ones, encode the same preferences. Finally, in <xref ref-type="sec" rid="c6-s7.s3">Section 6.7.3</xref>, we show qualitative <target target-type="page" id="pges_172"/>comparisons and demonstrate the usefulness of our approach in a downstream task consisting of &#x2018;search by similarity&#x2019;.</p>
<sec id="c6-s7.s1">
<label>6.7.1</label>
<title>Human Judgment Perceptual Similarity of Drape</title>
<p>We then use ten samples from our real dataset with known ground truth mechanical parameters and setup the user study as follows. Participants are presented with a triplet of fabrics and, using one fabric as a reference, they are asked which of the two remaining fabrics is most similar [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>; Gar+14] to the reference fabric. Participants are encouraged to manipulate the samples and focus only on the mechanical similarity and overall drape, and to ignore properties like material reflectance. Each participant rated 20 triplets that were pseudo-randomly sampled, ensuring that at least each of the ten test fabrics is used twice as a reference. Given the same triplet, we observe an average of 86.68% agreement between our participants across all experiments and materials, suggesting that there is a perceptual understanding of fabrics mechanics that humans share. We did not observe any significant differences in agreement depending on the volunteer demographics or level of expertise in fabric handling or simulation.</p>
<p>Leveraging the user study described above, we can compute an embedding that captures the relative distance between real materials according to human perception (i.e., ground truth perceptual similarity). <xref ref-type="fig" rid="fig-6.8">Figure 6.8</xref> depicts such embedding in 2D, computed using tSTE [<xref ref-type="bibr" rid="CIT382">Tam+11</xref>], which allows us to calculate perceptual distances between materials using the Euclidean norm. The embedding depicts many interesting patterns, including perfectly separated woven and knitted fabrics; elastic and thin fabrics are clustered together; thick and thin knits are well separated. These patterns suggest that fabric structure, composition, and density play an important, non-linear role in the overall perception of the mechanical properties of fabrics.</p>
<fig id="fig-6.8">
<label>Figure 6.8:</label>
<caption><title>Human Judgments 2D tSTE embedding [<xref ref-type="bibr" rid="CIT382">Tam+11</xref>] computed from human perceptual judgments about real fabric similarity (i.e., ground truth). We observe interesting patterns: woven and knits are separated; elastic materials are clustered; thick and thin materials are separated. Neither axes directly correspond to any material property, instead they emerge from the embedding.</title></caption>
<alt-text>Human Judgments 2D tSTE embedding [Tam+11] computed from human perceptual judgments about real fabric similarity (i.e., ground truth). We observe interesting patterns: woven and knits are separated; elastic materials are clustered; thick and thin materials are separated. Neither axes directly correspond to any material property, instead they emerge from the embedding.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.8.jpg"/>
</fig>
</sec>
<sec id="c6-s7.s2">
<label>6.7.2</label>
<title>Image-based vs. Human Perceptual Similarity</title>
<p>We evaluate the agreement between humans and our estimations within the context of a similarity rank. For each fabric of our real dataset, we compute the distance to the rest of the materials using different metrics: 1) the Euclidean distances on the Human Judgments tSTE embedding shown <xref ref-type="sec" rid="c6-s7.s1">Section 6.7.1</xref>; 2) the z-score distances in the parametric space of the mechanical simulation, <inline-formula><mml:math id="M439" display='block'><mml:mi mathvariant="script">P</mml:mi></mml:math></inline-formula>; 3) the distance using our similarity metric for drape explained in <xref ref-type="sec" rid="c6-s6">Section 6.6</xref>. We compare ground truth parameters and estimated ones for the second and third cases. The summary of results is shown in <xref ref-type="table" rid="c6-tab3">Table 6.3</xref>, and the complete analysis is presented in the supplementary material.</p>
<table-wrap id="c6-tab3">
<label>Table 6.3:</label>
<caption><title>Average (&#x00B1; std.) correlation with rankings obtained through the human judgments tSTE Embedding, depending on the parameter source (ground truth or predicted), and metric used to compute similarity (parameter distance or our drape similarity).</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center"><p><bold>Parameter Distance</bold></p></th>
<th valign="top" align="center"><p><bold> Drape Similarity Metric</bold></p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center"><bold>GT</bold></td>
<td valign="top" align="center"><p>0.431 &#x00B1; 0.17</p></td>
<td valign="top" align="center"><p><bold>0.893</bold>&#x00B1;0.05</p></td>
</tr>
<tr>
<td valign="top" align="center"><bold>Estimated</bold></td>
<td valign="top" align="center"><p>0.377 &#x00B1; 0.19</p></td>
<td valign="top" align="center"><p><bold>0.680</bold>&#x00B1;0.09</p></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><target target-type="page" id="pges_173"/>First, we demonstrate that our similarity metric using ground truth parameters correlates with human judgments. <xref ref-type="fig" rid="fig-6.9">Figure 6.9</xref> (top) illustrates the outcome using Spearman correlation (r). We observe that the correlation for each fabric is higher than 0.8, with an average of 0.893, showcasing a strong correlation. These results suggest that our metric, which only takes images of the <italic>hanging</italic> and <italic>stretch</italic> scenes as input, can measure distances between materials as humans would. Then, as shown in <xref ref-type="fig" rid="fig-6.9">Figure 6.9</xref> (bottom), we use our estimated parameters instead of the ground truth, reaching an average correlation of 0.680. Even though this value is slightly smaller than the ground truth, it is still significant to conclude the estimations of our model agree with human judgments. Note that the materials with higher correlation are those lying on the extreme areas of the embedding obtained in <xref ref-type="fig" rid="fig-6.8">Figure 6.8</xref> that have very distinct characteristics. Likewise, we also compute the distance for each material using merely the parameter spaces of the simulation. As can be seen, using this space does not produce correlated outputs with human judgments, reaching correlations below 0.44 in any case tested.</p>
<fig id="fig-6.9">
<label>Figure 6.9:</label>
<caption><title>Correlation between the ordering provided by the Human Judgments (x-axis) and our drape similarity metric with (top) the Ground Truth simulation parameters (y-axis), and (bottom) the estimations of our model. We plot z-scores instead of the raw distances to help visualization.</title></caption>
<alt-text>Correlation between the ordering provided by the Human Judgments (x-axis) and our drape similarity metric with (top) the Ground Truth simulation parameters (y-axis), and (bottom) the estimations of our model. We plot z-scores instead of the raw distances to help visualization.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.9.jpg"/>
</fig>
</sec>
<sec id="c6-s7.s3">
<label><target target-type="page" id="pges_174"/>6.7.3</label>
<title>Qualitative Results</title>
<p><xref ref-type="fig" rid="fig-6.10">Figure 6.10</xref> compares the simulations obtained with the parameters of our method with the ground truth parameters for a few examples. We can observe that our estimations are very close to the ground truth in this cases. Finally, our metric can be used to search between materials of similar drape. We illustrate this in <xref ref-type="fig" rid="fig-6.11">Figure 6.11</xref>. As shown, a naive ranking using directly the parameter space does not provide any meaningful ordering. On the contrary, using our similarity metric, we obtain ranks that agree with those given by the human embedding, showcasing the potential of our automatic metric to explore fabric collections.</p>
<fig id="fig-6.10">
<label>Figure 6.10:</label>
<caption><title>A comparison between the simulations obtained through the ground truth parameters, and those obtained using the predictions of our model, from a representative set of fabrics of our test set. As shown, the estimations of the model yield similar drapes to those of their ground truth counterparts, which we also evaluate quantitatively.</title></caption>
<alt-text>A comparison between the simulations obtained through the ground truth parameters, and those obtained using the predictions of our model, from a representative set of fabrics of our test set. As shown, the estimations of the model yield similar drapes to those of their ground truth counterparts, which we also evaluate quantitatively.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.10.jpg"/>
</fig>
<fig id="fig-6.11">
<label>Figure 6.11:</label>
<caption><title>Search by drape similarity. The ordering provided by the parameter space (first row) does not match human judgments (second row), while the arrangement obtained by our metric matches humans with high consensus (third row).</title></caption>
<alt-text>Search by drape similarity. The ordering provided by the parameter space (first row) does not match human judgments (second row), while the arrangement obtained by our metric matches humans with high consensus (third row).</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.11.jpg"/>
</fig>
<p><bold>Limitations</bold> Even if our model provides accurate predictions, its estimations are not always truthful to the real materials. We illustrate this in <xref ref-type="fig" rid="fig-6.12">Figure 6.12</xref>, where the model predicts fewer bends on the final drape than what the ground truth generates. According to our metric, this prediction is closer to thicker materials than to what humans perceive as most similar to the reference material.</p>
<fig id="fig-6.12">
<label>Figure 6.12:</label>
<caption><title>A failure case of our method. Given the reference material (first row) as input, the model predicts fewer bends than the ground truth. According to our metric, this prediction is closer to a thicker material (middle row), than to what humans perceive as most similar to the reference fabric (bottom row).</title></caption>
<alt-text>A failure case of our method. Given the reference material (first row) as input, the model predicts fewer bends than the ground truth. According to our metric, this prediction is closer to a thicker material (middle row), than to what humans perceive as most similar to the reference fabric (bottom row).</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.12.jpg"/>
</fig>
</sec>
</sec>
<sec id="c6-s8">
<label>6.8</label>
<title>Conclusions</title>
<p>In this chapter, we have presented a casual method to estimate the mechanical parameters of fabrics from depth images of fabric samples placed at two specific configurations. We have validated our architecture and inputs numerically, proving that all the components of our method are necessary to provide accurate estimations. While our quantitative analysis helped us understand the importance of each component, we found that these errors are not interpretable, nor do they help us understand the overall appearance of the predicted drape. Therefore, we have presented the first metric, which, by purely working on the image space, can capture differences in fabric mechanics as humans do. We have used such metric for two purposes: first, to validate the accuracy of our estimated parameters perceptually, and second, to showcase a novel application of search by drape similarity.</p>
<p><target target-type="page" id="pges_175"/>Our work could be improved in several ways. Our neural network is trained using a pure regression loss. Training the network using differentiable simulation could improve training, and help generalization and error interpretation. We could incorporate our perceptual metric as a loss function. However, it requires multiple differentiable simulations and a deep feature extractor, which will result in a significant computational overhead. In addition, less expressive</p>
<p>simulation engines may correlate less with human perception. Similarly, we would like to scale our training dataset and user study to handle more and more diverse samples, to cover a broader variety of fabric families. Interesting possible extensions would include taking as input RGB images instead of depth maps, training with real samples, incorporating symmetry consistency losses, or learning a similarity metric that can work by using captured images as input (instead of simulations). Finally, we hope our work might inspire future work in the long-standing problem of validating fabric mechanics in a way that is agnostic to the simulator parameters.<target target-type="page" id="pges_176"/><target target-type="page" id="pges_177"/></p>
</sec>
<sec id="c6-sA">
<label><target target-type="page" id="pges_178"/>6.A</label>
<title>Additional Implementation Details</title>
<p><bold>Training.</bold> We train our model using PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>] as our learning framework and Kornia [<xref ref-type="bibr" rid="CIT334">Rib+20</xref>] for image-processing operations, including normalization and data augmentation. The model is optimized using AdamW [<xref ref-type="bibr" rid="CIT255">LH17</xref>] for a total of 50 epochs, with a learning rate of &#x03B1; = 0.0003, which is halved every 10 iterations. All other optimizer hyperparameters followed the default AdamW configuration. We use mixed precision training [<xref ref-type="bibr" rid="CIT274">Mic+18</xref>], a batch size of 256, and a weight decay of 0.000002. The input images are of 180 &#x00D7; 180 pixels, and we use a Mean Squared Error (MSE, <italic>&#x2113;</italic><sub>2</sub>) as our loss function for our regression problem. Model size design and hyperparameter selection were conducted using Bayesian hyperparameter tuning [<xref ref-type="bibr" rid="CIT032">Bie20</xref>]. Training takes approximately one hour on an NVIDIA 2080 RTX GPU.</p>
<p><bold>Network Design.</bold> Our feature extractor is a pre-trained ResNet-18 [<xref ref-type="bibr" rid="CIT139">He+16</xref>] (TorchVision pretrained weights checkpoint weights=&#x2019;IMAGENET1K V1&#x2019;), followed by a Self-Attention layer [<xref ref-type="bibr" rid="CIT443">Zha+19b</xref>] with 64 attention heads. The output of this layer is max and average-pooled, and a final multi-layer perceptron (MLP) with 4 hidden layers with 512 hidden units each, receives the pooled features and the density value, and outputs the 6 predicted mechanical parameters. The MLP layers use ReLU [<xref ref-type="bibr" rid="CIT260">MHN13</xref>] non-linearities, followed by Layer Normalization [<xref ref-type="bibr" rid="CIT018">BKH16</xref>], and a Dropout [<xref ref-type="bibr" rid="CIT373">Sri+14</xref>] rate of 0.5. Both the feature extractor and the MLP are trained simultaneously.</p>
<p><bold>Data Augmentation.</bold> We use the following policy, in order: First, we randomly rescale the input images using (0.5, 1.5) rescale ranges. We then randomly erase [<xref ref-type="bibr" rid="CIT452">Zho+20</xref>] parts of the input images (<italic>p</italic> = 0.2), randomly rotate them (<italic>p</italic> = 0.2, <italic>angle</italic> = &#x00B1;10<sup>&#x25E6;</sup>), apply a random perspective change (<italic>p</italic> = 0.5, <italic>scale</italic> = 0.5), and a random thin plate spline warp (<italic>p</italic> = 0.2, <italic>scale</italic> = 0.5). We then crop the center area, with a resolution of 180 &#x00D7; 180 pixels. To those images, we randomly change their contrast and brightness, on the (0.5, 1.5) ranges. We then apply a final set of transformations, including image equalization, random horizontal flip (<italic>p</italic> = 0.5), random Gaussian noise (<italic>p</italic> = 0.5, <italic>&#x03BC;</italic> = 0, &#x03C3; = 0.02), random posterization ((<italic>p</italic> = 0.5)), random sharpening (<italic>p</italic> = 0.5), and random Gaussian blur ((<italic>p</italic> = 0.5), kernel size of 3 &#x00D7; 3). Please refer to the Kornia documentation [<xref ref-type="bibr" rid="CIT334">Rib+20</xref>] for specific implementation details of each of these transformations and the specific meaning of the mentioned parameters.</p>
</sec>
<sec id="c6-sB">
<label>6.B</label>
<title>Additional Results</title>
<p>In this section, we show the relative similarity rankings computed through the tSTE embedding of our user study, compared to the ground truth and predicted parameter distances, as well as the drape similarity computed using the <target target-type="page" id="pges_179"/>ground truth and predicted parameters. We also provide additional details of our user study.</p>
<fig id="fig-6.13">
<label>Figure 6.13:</label>
<caption><title>Relationship between the distance obtained by our user study, our distance metric, and the difference in each parameter. On the top row, we plot the difference between the value of a parameter against the average distance between fabrics we obtained through our user study. On the bottom, we show the same parameter distances, now against the distance predicted through our perceptual metric. As shown, no mechanical parameter dominates the distance predicted by either the user study or our metric. The perception of mechanical properties can be understood as a highly non-linear phenomenon in which many parameters interact in complex ways. We use z-scores to help visualization.</title></caption>
<alt-text>Relationship between the distance obtained by our user study, our distance metric, and the difference in each parameter. On the top row, we plot the difference between the value of a parameter against the average distance between fabrics we obtained through our user study. On the bottom, we show the same parameter distances, now against the distance predicted through our perceptual metric. As shown, no mechanical parameter dominates the distance predicted by either the user study or our metric. The perception of mechanical properties can be understood as a highly non-linear phenomenon in which many parameters interact in complex ways. We use z-scores to help visualization.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.13.jpg"/>
</fig>
<fig id="fig-6.14">
<label>Figure 6.14:</label>
<caption><title>To recreate possible misalignments that may be present on the real images, we propose a <italic>simulation-space</italic> data augmentation policy, in which we simulate the drapes from different rest positions, on the [&#x2212;5, 5] degree ranges. For every scene, we generate 11 different simulations uniformely in this range, effectively creating a dataset of images for a single material.</title></caption>
<alt-text>To recreate possible misalignments that may be present on the real images, we propose a simulation-space data augmentation policy, in which we simulate the drapes from different rest positions, on the [&#x2212;5, 5] degree ranges. For every scene, we generate 11 different simulations uniformely in this range, effectively creating a dataset of images for a single material.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.14.jpg"/>
</fig>
<sec id="c6-sB.1">
<title>User Study: Procedure and Stimuli</title>
<p><bold>Stimuli.</bold> We use ten physical fabric samples of our real dataset, which are representative of different types of fabric structures, thicknesses, densities, and mechanical properties. We leverage 25&#x00D7;25 cm samples, which is an adequate size for the manipulation of the materials.</p>
<fig id="fig-6.15">
<label><target target-type="page" id="pges_180"/>Figure 6.15:</label>
<caption><title>Ranking classification accuracy on different ways of comparing materials. Comparing materials simply by measuring their distances on parameter space yields innacurate similarity rankings, as the parameter space is not orthogonal and cannot be linearly correlated with human perception of material similarity. Our Drape similarity metric provides us with a way of assessing the similarity between materials on a more perceptual-aware space, yielding more accurate rankings.</title></caption>
<alt-text>Ranking classification accuracy on different ways of comparing materials. Comparing materials simply by measuring their distances on parameter space yields innacurate similarity rankings, as the parameter space is not orthogonal and cannot be linearly correlated with human perception of material similarity. Our Drape similarity metric provides us with a way of assessing the similarity between materials on a more perceptual-aware space, yielding more accurate rankings.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.15.jpg"/>
</fig>
<p><bold>Participants.</bold> A total of 30 volunteers took part in the test. The participants had different backgrounds and diverse levels of expertise in physical manipulation of textiles. A study on their demographics can be found in <xref ref-type="fig" rid="fig-6.20">Figure 6.20</xref>.</p>
<p><bold>Procedure.</bold> Following previous work on human perception [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>; Gar+14], we design a Two-Alternative Forced Choice (2AFC) user study in which participants are presented with triplets of physical fabric samples. We performed the experiment in a lab-controlled environment because the fabric samples had to be physically manipulated. Participants were asked to select which of two fabrics is most similar to a reference fabric. They were suggested to manipulate the samples (stretching, bending them, etc.) and to focus only on mechanical similarity, thus ignoring other factors such as optical properties (color, specularity, transparency) or irrelevant tactile feedback (softness, for instance). Material reflectance was thus ignored by participants. Volunteers were not instructed to leverage fabric motion to make their decisions but were free to dynamically manipulate the samples, which some participants did. Following recommendations in user study design for graphics [<xref ref-type="bibr" rid="CIT039">Byl+22</xref>; <xref ref-type="bibr" rid="CIT074">Dod+22</xref>], we informed the users that there are no right or wrong answers, allowed them to use their own criteria for answering the questions, and self-identify their demographics and levels of expertise. The volunteer demographics can be found on 6.20.</p>
<p><target target-type="page" id="pges_181"/>Each participant rated 20 triplets that were pseudo-randomly sampled following some rules: each of the 10 fabrics in our test set is shown as a reference at least twice to each participant, the order of the triplets is random, and each triplet was evaluated by at least 2 participants. Each test took between 15 to 30 minutes, depending on the participant, and was followed by an open discussion, to better understand which criteria each participant used to make their decisions. We started each test with a brief interview to register the self-reported participants&#x2019; demographics [<xref ref-type="bibr" rid="CIT074">Dod+22</xref>], we continued it with the 20 perceptual comparisons and finished it with an open discussion so as to better understand which factors each participant used to make their decisions.</p>
<fig id="fig-6.16">
<label>Figure 6.16:</label>
<caption><title>Correlation between the ordering provided by the Human Judgments (x-axis) and our Drape Similarity Metric (y-axis). We plot z-scores instead of the raw distances to help visualization.</title></caption>
<alt-text>Correlation between the ordering provided by the Human Judgments (x-axis) and our Drape Similarity Metric (y-axis). We plot z-scores instead of the raw distances to help visualization.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.16.jpg"/>
</fig>
<p><bold>Results: Agreement</bold> 30 participants, with diverse levels of expertise with fabrics and computer simulations, volunteered for this user study. Given the same triplet, we observe an average of 86.68% agreement between our participants across all experiments and materials, suggesting that there is a perceptual understanding of fabrics mechanics that humans share. We did not observe any significant differences in agreement depending on the volunteer demographics, or level of expertise on fabric handling or simulation. This suggests that the perception of similarity of the mechanical properties of textiles may be relatively universal. When prompted, participants generally admitted that some triplets were easier to evaluate than others, either because neither candidate fabrics were similar to the reference or because both were very similar and thus difficult to make a definitive decision. Participants tended to evaluate elasticity and the shape of the fabric when it bends. When in doubt about those factors, a standard procedure was to assess similarity using the perceived weight of the material.</p>
<fig id="fig-6.17">
<label><target target-type="page" id="pges_182"/>Figure 6.17:</label>
<caption><title>Correlation between the ordering provided by our similarity metric with the Ground Truth parameters (x-axis) and the estimated ones (y-axis). We plot z-scores instead of the raw distances to help visualization.</title></caption>
<alt-text>Correlation between the ordering provided by our similarity metric with the Ground Truth parameters (x-axis) and the estimated ones (y-axis). We plot z-scores instead of the raw distances to help visualization.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.17.jpg"/>
</fig>
</sec>
</sec>
<sec id="c6-sC">
<label>6.C</label>
<title>Ablation Study: Similarity Metric Parameters</title>
<p>Here, we study the factors that might impact the performance of the metric: the image-based similarity metric (<italic>IM</italic>), the number of simulations necessary to account for the non-determinism of the simulation (<italic>N</italic> ), and the scene (<italic>hanging</italic> or <italic>stretch</italic>).</p>
<list list-type="bullet">
<list-item><p><bold>Choice of</bold> <italic>IM</italic> The image-based metric should be able to capture the subtle differences in wrinkles, folds, and overall shape of the simulated drape. While previous work has learned this metric from data [<xref ref-type="bibr" rid="CIT102">Gar+14</xref>; <xref ref-type="bibr" rid="CIT225">Lag+19</xref>], we proposed a simpler approach using existing image-based similarity metrics. <xref ref-type="fig" rid="fig-6.21">Figure 6.21 (a)</xref> shows results for several deep and non-deep methods. For deep learning-based metrics, it is particularly relevant that the metric can account for local and positional differences. A <italic>Content loss</italic> or the <italic>LPIPS</italic> shows stronger performance than translation-invariant losses (e.g. <italic>Style Loss</italic> or <italic>DISTS</italic>).</p></list-item>
<list-item><p><bold>Choice of</bold> <italic>N</italic> We study how many simulations that are needed to account for the non-determinism of the simulation. <xref ref-type="fig" rid="fig-6.21">Figure 6.21 (b)</xref> shows that <italic>N</italic> &#x2265; 5 is enough to obtain accurate results.</p></list-item>
<list-item><p><bold>Choice of scene</bold> We study which scene is best for estimating drape similarity in <xref ref-type="fig" rid="fig-6.21">Figure 6.21 (c)</xref>. The <italic>stretch</italic> scene generally provides more information than the <italic>hanging</italic> scene, while using both achieves the best agreement overall.</p></list-item>
</list>
<fig id="fig-6.18">
<label><target target-type="page" id="pges_183"/>Figure 6.18:</label>
<caption><title>Correlation between the ordering provided by the tSTE embedding from our user study (x-axis) and our distance metric, computed using the estimated parameters (y-axis). We plot z-scores instead of the raw distances to help visualization.</title></caption>
<alt-text>Correlation between the ordering provided by the tSTE embedding from our user study (x-axis) and our distance metric, computed using the estimated parameters (y-axis). We plot z-scores instead of the raw distances to help visualization.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.18.jpg"/>
</fig>
<fig id="fig-6.19">
<label>Figure 6.19:</label>
<caption><title>Correlation between the ordering provided by the tSTE embedding from our user study (x-axis) and the distance computed on the parameter space, for the model predictions (y-axis). We plot z-scores instead of the raw distances to help visualization.</title></caption>
<alt-text>Correlation between the ordering provided by the tSTE embedding from our user study (x-axis) and the distance computed on the parameter space, for the model predictions (y-axis). We plot z-scores instead of the raw distances to help visualization.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.19.jpg"/>
</fig>
<fig id="fig-6.20">
<label><target target-type="page" id="pges_184"/>Figure 6.20:</label>
<caption><title>Demographics of the participants of our perceptual study. Following recommendations on user studies for computer graphics [<xref ref-type="bibr" rid="CIT074">Dod+22</xref>; <xref ref-type="bibr" rid="CIT039">Byl+22</xref>], we capture the self-identified gender, age, and education level of each participant. We also asked them to identify their level of expertise with handling physical textile materials, as well as their expertise in computer simulation and rendering of fabrics.</title></caption>
<alt-text>Demographics of the participants of our perceptual study. Following recommendations on user studies for computer graphics [Dod+22; Byl+22], we capture the self-identified gender, age, and education level of each participant. We also asked them to identify their level of expertise with handling physical textile materials, as well as their expertise in computer simulation and rendering of fabrics.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.20.jpg"/>
</fig>
<fig id="fig-6.21">
<label><target target-type="page" id="pges_185"/>Figure 6.21:</label>
<caption><title>Ablation study on different factors of our similarity metric. Y-axis is Spearman correlation. (a) Results for several image-based metrics <italic>IM</italic> (we use the implementations provided in <italic>PIQ</italic> [<xref ref-type="bibr" rid="CIT199">KZP19</xref>]). (b) Number of simulations needed to account for the non-determinism (optimal choice is 5). (c) Impact of the scene.</title></caption>
<alt-text>Ablation study on different factors of our similarity metric. Y-axis is Spearman correlation. (a) Results for several image-based metrics IM (we use the implementations provided in PIQ [KZP19]). (b) Number of simulations needed to account for the non-determinism (optimal choice is 5). (c) Impact of the scene.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-6.21.jpg"/>
</fig>
</sec>
</body>
<back>
<fn-group>
<fn id="c6-fn1"><label>1</label> <p>Additional publication details and results are included on <bold>the project website</bold>.</p></fn>
</fn-group>
</back>
</book-part>
<book-part id="c7" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_186"/><target target-type="page" id="pges_187"/>Chapter 7.</label>
<title>Neural Models for Global Illumination</title>
</title-group>
</book-part-meta>
<body>
<fig id="fig-7.1">
<label>Figure 7.1:</label>
<caption><title>Renders using analytical illumination methods with multiple importance sampling and our learned lighting approach for two environment maps.</title></caption>
<alt-text>Renders using analytical illumination methods with multiple importance sampling and our learned lighting approach for two environment maps.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.1.jpg"/>
</fig>
<p>Environment maps are commonly used to represent and compute far-field illumination in virtual scenes. However, they are expensive to evaluate and sample from, limiting their applicability to real-time rendering. Previous methods have focused on compression through spherical-domain approximations, or on learning priors for natural, day-light illumination. These hinder both accuracy and generality, and do not provide the probability information required for importance-sampling Monte Carlo integration. In this chapter, we propose NEnv, a deep-learning fully-differentiable method, capable of compressing and learning to sample from a single environment map. NEnv is composed of two different neural networks: A normalizing flow, able to map samples from uniform distributions to directions of the environment map, also providing their corresponding probabilities; and an implicit neural representation which compresses the environment map into an efficient differentiable function. The computation time of environment samples with NEnv is two orders of magnitude less than with traditional methods. NEnv makes no assumptions regarding environment map content, achieving higher generality than previous learning-based approaches. The contributions in this chapter led to the following publication, currently under review:</p>
<fig id="fig-14">
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-14.jpg"/>
</fig>
<disp-quote>
<p><target target-type="page" id="pges_188"/>&#x201C;NEnv: Neural Environment Maps for Global Illumination&#x201D;</p>
<p>Carlos Rodriguez-Pardo, Javier Fabre, Elena Garces, Jorge Lopez-Moreno</p>
<p><italic>Under Review</italic> (2023)</p>
</disp-quote>
<fig id="fig-7.2">
<label>Figure 7.2:</label>
<caption><title>From left to right, Ground Truth and our learned RGB environment map compression and sampling algorithms, and renders using both illumination configurations on a complex scene.</title></caption>
<alt-text>From left to right, Ground Truth and our learned RGB environment map compression and sampling algorithms, and renders using both illumination configurations on a complex scene.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.2.jpg"/>
</fig>
<sec id="c7-s1">
<label>7.1</label>
<title>Introduction</title>
<p>Environment Maps are widely used in rendering to represent far-field illumination around a point with a single texture, usually a High Dynamic Range image (HDRi). There are two typical scenarios for using these maps: as infinite spherical light sources in offline rendering methods (such as path tracing), or as local probes that encode near-field irradiance to approximate global illumination in real-time applications.</p>
<p><target target-type="page" id="pges_189"/>In the first case, many algorithms such as Multiple Importance Sampling (MIS) [<xref ref-type="bibr" rid="CIT399">VG95</xref>] require the ability to sample lights in the scene, which can be problematic when dealing with environment maps. Since the source data comes from an image, there is no analytic Cumulative Distribution Function (CDF) that can be used for sampling. Instead, tabulated methods can be used, but these require a pre-computation step that consumes a considerable amount of memory, as well as a search step to find the sample probability.</p>
<p>In real-time scenarios, irradiance sampling is becoming more relevant with the increasing capabilities of GPUs and sample reservoir techniques for global illumination [<xref ref-type="bibr" rid="CIT033">Bit+20</xref>; <xref ref-type="bibr" rid="CIT299">Ouy+21</xref>], yet the usual techniques to sample environment maps do not easily benefit from parallel GPU environments. Also, in real-time, analytical approximations (such as Spherical Harmonics or Spherical Gaussians) are often used to overcome this issue, however, they are not able to capture final-scale details present in most environment maps.</p>
<p>This work presents a novel learning-based method for sampling, PDF evaluation, and compression of high-resolution environment maps, able to encode any type of illumination with high fidelity. Our method builds on recent advances in neural representations, that we found particularly suitable for this kind of problem, where it is important to encode both the high-frequency and low-frequency patterns. Key to our solution is to use a normalizing flow [<xref ref-type="bibr" rid="CIT080">Dur+19a</xref>] to encode an <italic>invertible</italic> representation of the environment map PDF, reducing both sampling and evaluation time by two orders of magnitude compared with analytic solutions. Thanks to being invertible, our representation is differentiable and compatible with importance-sampling techniques. We also propose a method based on implicit neural representations to compress the environment map, reducing the memory footprint of the original image with minimal loss. We demonstrate that our approach works for a wide range of scenes and is faster and more accurate than previous methods.</p>
</sec>
<sec id="c7-s2">
<label>7.2</label>
<title>Related Work</title>
<p>In offline rendering, environment maps are usually encoded in high resolution (&#x003E; 4<italic>K</italic>) HDR images (Radiance HDR [<xref ref-type="bibr" rid="CIT228">LS98</xref>], OpenEXR [<xref ref-type="bibr" rid="CIT189">KBH03</xref>]), requiring tabulated approximations for sampling and evaluation. Techniques such as Piecewise&#x015B;Constant 2D Distributions [<xref ref-type="bibr" rid="CIT308">PJH16</xref>] or Hierarchical Warping [<xref ref-type="bibr" rid="CIT053">Cla+21</xref>] are used for this purpose, but they are not suited for RT applications due to limited time and memory budget, being a costly and difficult-to-parallelize operation, even for offline rendering. This cost can be reduced, for instance, by taking into account scene information, such as occluders [<xref ref-type="bibr" rid="CIT002">Aga+03</xref>] to avoid poor-quality samples.</p>
<p><target target-type="page" id="pges_190"/>Heitz [<xref ref-type="bibr" rid="CIT143">Hei20</xref>] proposed a method to invert non&#x015B;analytically invertible CDFs, which can be applied to some distribution functions (e.g.: BSDFs) to obtain an analytical sampling function. However, environment maps are not good candidates for triangle-cut parametrization, since their PDF and CDF come from a discrete tabulated source (HDR image).</p>
<p>To avoid using high resolution images, real time methods rely on environment map prefiltering [<xref ref-type="bibr" rid="CIT201">Kau+00</xref>], or analytical approximations, such as Spherical Harmonics (SH) [<xref ref-type="bibr" rid="CIT356">See66</xref>], and Spherical Gaussians (SG) [<xref ref-type="bibr" rid="CIT429">Xu+13</xref>]. These methods allow for analytical evaluation of an environment map, as well as analytical sampling (directly, in the case of SG, or using Hierarchical Sampling for SH [<xref ref-type="bibr" rid="CIT181">JCJ09</xref>]) and easy interpolation (useful when not enough samples can be requested). Still, they do not accurately represent complex HDR images commonly used when representing high&#x015B;detailed light setups by using environment maps, and blending important light features in the original images, sometimes generating artifacts. More recently, neural representations have been explored for this purpose. For instance, Gardner <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT105">GES22</xref>] propose a conditional neural field representation, based on a variational auto-decoder (RENI), which leverages natural image priors to efficiently encode a full HDR environment image into</p>
<p>a few-dimensional latent space vector. Being low-dimensional, it shares the same accuracy limitations that its predecessors, and does not consider sampling, or evaluation in its design. In the following, we describe the most relevant neural approaches for sampling and 2D image representations.</p>
<sec id="c7-s2.s1">
<label>7.2.1</label>
<title>Neural Sampling and Representations</title>
<p><bold>Learned Sampling and Normalizing Flows</bold> Normalizing Flows (NFs) have been proposed as powerful models for learned sampling for rendering applications. The seminal work of M&#x00FC;ller <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT284">M&#x00FC;l+19</xref>] proposed <italic>Neural Importance Sampling</italic>, using <italic>Piecewise-Linear</italic> and <italic>Piecewise-Quadratic Coupling Flows</italic> for learned sampling for Monte Carlo rendering. This method showed the first successful use of these neural models for Path-Tracing, but their results showed a prohibitive overhead in runtime cost over previous methods. In a later work, M&#x00FC;ller <italic>et al.</italic>proposed <italic>Neural Control Variates</italic> [<xref ref-type="bibr" rid="CIT285">M&#x00FC;l+20</xref>], an <italic>Autorregresive Flow</italic> which allows for more efficient unbiased integration. Beyond Monte Carlo integration, Flows have been used for BRDF representations in inverse rendering [<xref ref-type="bibr" rid="CIT050">CNN22</xref>]. Relatedly, Sztrajman <italic>et al.</italic> [<xref ref-type="bibr" rid="CIT380">Szt+21</xref>] propose a method for neural sampling of BRDFs. However, instead of relying on NFs, they propose an encoder network that maps from <italic>Neural BRDF</italic> to fitted Blinn-Phong parameters for which importance sampling is known. Besides these tasks, NFs have been utilized in computer graphics and vision for image [<xref ref-type="bibr" rid="CIT207">KD18</xref>; <xref ref-type="bibr" rid="CIT406">WZY22</xref>] and video [<xref ref-type="bibr" rid="CIT217">Kum+19</xref>] generation, compression [<xref ref-type="bibr" rid="CIT146">Hel+20</xref>], super-resolution [<xref ref-type="bibr" rid="CIT257">Lug+20</xref>], domain translation [<xref ref-type="bibr" rid="CIT119">Gro+20</xref>], and uncertainty quantification [<xref ref-type="bibr" rid="CIT412">Wan+22c</xref>; <xref ref-type="bibr" rid="CIT357">Sel+20</xref>]. We build upon <italic>Neural Importance Sampling</italic> <target target-type="page" id="pges_191"/>[<xref ref-type="bibr" rid="CIT284">M&#x00FC;l+19</xref>] and propose a lightweight model for sampling and PDF evaluation of environment maps, as we describe in <xref ref-type="sec" rid="c7-s4">Section 7.4</xref>. We show that the proposed architecture and training produce a single network that can be integrated into any rendering pipeline, providing a significant speedup in image-based lighting with importance sampling.</p>
<p><bold>Implicit Neural Representations</bold> (INRs) have emerged in recent years as a powerful alternative to traditional representations for natural signals. They allow for differentiable, continuous, and expressive functions which can represent many types of data, such as images [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>; <xref ref-type="bibr" rid="CIT267">Meh+21</xref>], video [<xref ref-type="bibr" rid="CIT099">Gao+21</xref>], or <italic>Neural Fields</italic> for rendering [<xref ref-type="bibr" rid="CIT276">Mil+20</xref>]. INRs typically build upon standard Multi Layer Perceptrons (MLPs) but benefit from working on the frequency domain to better handle higher frequencies, which are present on natural signals. These can be introduced to the models using fixed <italic>Positional Encodings</italic> [<xref ref-type="bibr" rid="CIT276">Mil+20</xref>; <xref ref-type="bibr" rid="CIT024">Bar+21</xref>], <italic>Fourier Features</italic> [<xref ref-type="bibr" rid="CIT384">Tan+20</xref>; <xref ref-type="bibr" rid="CIT195">Kar+21</xref>] or <italic>Sinusoidal Activations</italic> [<xref ref-type="bibr" rid="CIT105">GES22</xref>; <xref ref-type="bibr" rid="CIT267">Meh+21</xref>; <xref ref-type="bibr" rid="CIT366">Sit+20</xref>], among other approaches. These general-purpose models have been utilized extensively in inverse and neural rendering, for representing geometry [<xref ref-type="bibr" rid="CIT381">Tak+21</xref>; <xref ref-type="bibr" rid="CIT430">Yan+21</xref>], illumination [<xref ref-type="bibr" rid="CIT105">GES22</xref>; <xref ref-type="bibr" rid="CIT101">GMX22</xref>; <xref ref-type="bibr" rid="CIT015">Att+22</xref>], material reflectance [<xref ref-type="bibr" rid="CIT019">Baa+22</xref>; <xref ref-type="bibr" rid="CIT088">Fan+22</xref>; <xref ref-type="bibr" rid="CIT380">Szt+21</xref>; <xref ref-type="bibr" rid="CIT221">Kuz+22</xref>; <xref ref-type="bibr" rid="CIT220">Kuz+21</xref>; <xref ref-type="bibr" rid="CIT434">Yao+22</xref>] or entire scenes [<xref ref-type="bibr" rid="CIT276">Mil+20</xref>; <xref ref-type="bibr" rid="CIT094">Fri+22</xref>; <xref ref-type="bibr" rid="CIT383">Tan+22</xref>; <xref ref-type="bibr" rid="CIT401">Ver+22</xref>; <xref ref-type="bibr" rid="CIT367">Sit+21</xref>]. We refer the reader to the seminal surveys in <italic>Neural Fields</italic> [<xref ref-type="bibr" rid="CIT428">Xie+22</xref>; <xref ref-type="bibr" rid="CIT385">Tew+22</xref>] for more comprehensive reviews of these representations for neural rendering and computer vision. Besides their effectiveness for rendering, INRs have also been used for image [<xref ref-type="bibr" rid="CIT377">Str+22</xref>] and video [<xref ref-type="bibr" rid="CIT333">Rho+22</xref>] compression. We build upon these findings to design an efficient yet expressive INR for environment maps, as described in <xref ref-type="sec" rid="c7-s5">Section 7.5</xref>.</p>
</sec>
</sec>
<sec id="c7-s3">
<label><target target-type="page" id="pges_192"/>7.3</label>
<title>Overview</title>
<p>Our method takes as input a single HDRi environment map of any resolution, which is an image-based 360&#x00BA; representation of the illumination of a scene <inline-formula><mml:math id="M442" display='block'><mml:mi mathvariant="script">I</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. First, we perform a series of pre-processing steps (<xref ref-type="sec" rid="c7-s3.s1">Section 7.3.1</xref>) and compute the ground truth probability distribution function,(<xref ref-type="sec" rid="c7-s3.s2">Section 7.3.2</xref>). Then, we train a normalizing flow <inline-formula><mml:math id="M443" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> for joint sampling and PDF evaluation (<xref ref-type="sec" rid="c7-s4">Section 7.4</xref>), and an implicit neural representation <inline-formula><mml:math id="M444" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> for environment map compression (<xref ref-type="sec" rid="c7-s5">Section 7.5</xref>). We illustrate NEnv in <xref ref-type="fig" rid="fig-7.3">Figure 7.3</xref>.</p>
<fig id="fig-7.3">
<label>Figure 7.3:</label>
<caption><title>Overview of <italic>NEnv</italic>. We propose an invertible generative neural network <inline-formula><mml:math id="M440" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> (i.e. a <italic>normalizing flow</italic>) which simultaneously learns to sample directions (<italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>) from an environment map, as well as evaluating the probability of a given direction. This model, illustrated on the left, can be integrated seamlessly into path tracer render engines, allowing for efficient multiple importance sampling. We additionally use an implicit neural representation <inline-formula><mml:math id="M441" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> (second column) which learns to map from directions to linear RGB values, allowing for faster evaluation.</title></caption>
<alt-text>Overview of NEnv. We propose an invertible generative neural network F (i.e. a normalizing flow) which simultaneously learns to sample directions (&#x03B8;, &#x03C6;) from an environment map, as well as evaluating the probability of a given direction. This model, illustrated on the left, can be integrated seamlessly into path tracer render engines, allowing for efficient multiple importance sampling. We additionally use an implicit neural representation c (second column) which learns to map from directions to linear RGB values, allowing for faster evaluation.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.3.jpg"/>
</fig>
<p><inline-formula><mml:math id="M445" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> is an invertible neural network which can generate directional samples <inline-formula><mml:math id="M446" display='block'><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="script">F</mml:mi><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula> by mapping from a known and simple noise distribution <inline-formula><mml:math id="M447" display='block'><mml:mi>z</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>,</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x007E;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>z</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula>, as well as evaluating the probability density of any direction <italic>p</italic>(<italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>). <inline-formula><mml:math id="M448" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> provides an efficient and differentiable approximation of the PDF encoded in <inline-formula><mml:math id="M449" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula>, which is particularly useful for Monte Carlo integration in rendering systems using Multiple Importance Sampling (MIS). We model <inline-formula><mml:math id="M450" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> as a sinusoidal Multilayer Perceptron (MLP) which maps directions (<italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>) to RGB values: <italic>RGB</italic> = <italic>C</italic>(<italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>); compressing <inline-formula><mml:math id="M452" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula> into a differentiable, continuous function. We train <inline-formula><mml:math id="M453" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M454" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> individually for each input <inline-formula><mml:math id="M455" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula>, and design their architectures and training configurations to achieve high compression rates, efficient evaluation, and minimal loss in rendering quality. Once trained, both <inline-formula><mml:math id="M456" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M457" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> can be integrated seamlessly into rendering engines, considerably reducing the time and memory consumption per sample compared to traditional global illumination approaches.</p>
<p><target target-type="page" id="pges_193"/>Implementation details are included in <xref ref-type="sec" rid="c7-s6">Section 7.6</xref>. To validate our method, we perform extensive ablation studies, and perform comparisons with previous work on environment map approximations (<xref ref-type="sec" rid="c7-s8">Section 7.8</xref>), both in environment map reconstruction and its impact on final renders.</p>
<sec id="c7-s3.s1">
<label>7.3.1</label>
<title>Map Pre-processing</title>
<p>Some environment maps have large and intense light sources close to the horizontal borders of the image, separating the same light source into extreme values of <italic>&#x03C6;</italic> &#x2248; 0 and <italic>&#x03C6;</italic> &#x2248; 2<italic>&#x03C0;</italic>. While this is not a problem when rendering with traditional approaches, we observe these cases become very challenging for our normalizing flows, creating artifacts and tails in the encoded PDFs, as we show in <xref ref-type="fig" rid="fig-7.4">Figure 7.4</xref>. We solve this by rotating the environment map by <italic>&#x03C0;</italic> &#x2212; <italic>&#x03C6;<sub>max</sub></italic> angles, so that its most intense light source is in <italic>&#x03C6;</italic> = <italic>&#x03C0;</italic>, thus avoiding strong discontinuities at the borders:</p>
<p><disp-formula id="Eq007-1"><label>(7.1)</label> <mml:math id="M458" display='block'><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mo>max</mml:mo></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo>max</mml:mo></mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:munder><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:munder><mml:mi mathvariant="script">I</mml:mi><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo></mml:math></disp-formula></p>
<fig id="fig-7.4">
<label>Figure 7.4:</label>
<caption><title>Environment map and the PDF encoded by NEnv for an unprocessed (left) and rotated (right) environment map. Our simple rotation preprocessing avoids discontinuities in the borders of the learned reconstructions, enhancing the performance of our models.</title></caption>
<alt-text>Environment map and the PDF encoded by NEnv for an unprocessed (left) and rotated (right) environment map. Our simple rotation preprocessing avoids discontinuities in the borders of the learned reconstructions, enhancing the performance of our models.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.4.jpg"/>
</fig>
<p>While there are specific parameterizations for normalizing flows on spheres [<xref ref-type="bibr" rid="CIT331">Rez+20</xref>], we empirically observe that this rotation algorithm effectively solves the discontinuity issues we observed on real-world environment maps without requiring any modifications in our model design.</p>
</sec>
<sec id="c7-s3.s2">
<label>7.3.2</label>
<title>Computing Ground Truth Tabular PDFs</title>
<p>As mentioned, there are two widely&#x015B;used approaches for sampling directions according to an environment map distribution:</p>
<list list-type="bullet">
<list-item><p>Computing a <italic>Piecewise-Constant</italic> [<xref ref-type="bibr" rid="CIT308">PJH16</xref>] 2D Distribution by sampling a marginal PDF to select a row/column of <inline-formula><mml:math id="M459" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula>, then sampling a conditional PDF to choose an individual pixel.</p></list-item>
<list-item><p>A Hierarchical Warping [<xref ref-type="bibr" rid="CIT053">Cla+21</xref>; <xref ref-type="bibr" rid="CIT306">Pha19</xref>] algorithm based on MIP&#x015B;mapping over <inline-formula><mml:math id="M460" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula> to iteratively warp 2D samples from an uniform space until they match the target distribution.</p></list-item>
</list>
<p>Pharr [<xref ref-type="bibr" rid="CIT306">Pha19</xref>] compared these methods, stating that despite obtaining different spatial warpings, both produce similar error in renders. Interestingly, the official implementation for Mitsuba3 [<xref ref-type="bibr" rid="CIT177">Jak+22</xref>] relies on the latter approach, while the available implementation for the 4th edition of PBRT [<xref ref-type="bibr" rid="CIT307">PJH20</xref>] uses the former. These are both widely&#x015B;used engines in the literature. Due to its comparably better efficiency (see <xref ref-type="sec" rid="c7-s8.s3">Section 7.8.3</xref>), which is important during model training, we choose to use the Piecewise-Constant method to compute our ground truth PDFs and samples.</p>
<p><target target-type="page" id="pges_194"/>To obtain the PDF for a given environment map <inline-formula><mml:math id="M461" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula>, we first define <italic>f</italic>(<italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>), by mapping directions to the luminance values stored in the image <inline-formula><mml:math id="M462" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula>. These values are proportional to the probability of selecting a particular pixel that we want to compute, but require additional processing and normalization. We start by multiplying <inline-formula><mml:math id="M463" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula> by sin(<italic>&#x03B8;</italic>), given the <italic>&#x03B8;</italic> corresponding to each row, thus eliminating the distortion caused by mapping the image to the unit sphere [<xref ref-type="bibr" rid="CIT308">PJH16</xref>]. Using <italic>f</italic>(<italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>) as source data, we then compute one conditional density distribution <italic>f<sub>conditional&#x03B8;</sub></italic> (<italic>&#x03C6;</italic>) for each <italic>&#x03C6;</italic>, so we can later obtain a single marginal distribution by integrating all conditional distributions (rows) to build 1D Piecewise distributions. We use this marginal distribution to sample a row, and in turn, sample the corresponding conditional distribution stored for such row to select a column, hence getting a single pixel defined by <italic>&#x03B8;</italic>, <italic>&#x03C6;</italic> coordinates.</p>
<p>Finally, to compute the probability density of choosing a given pixel <inline-formula><mml:math id="M464" display='block'><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="script">I</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, we obtain the CDF for the conditional distribution at row <italic>&#x03C6;</italic> for value <italic>&#x03B8;</italic>, and normalize it to the marginal distribution integral:</p>
<p><disp-formula id="Eq007-2"><label>(7.2)</label> <mml:math id="M465" display='block'><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="script">I</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:msub><mml:mtext>conditional</mml:mtext><mml:mi>&#x03B8;</mml:mi></mml:msub></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub></mml:munderover><mml:msub><mml:mi>f</mml:mi><mml:mtext>marginal</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>where <italic>f<sub>conditional</sub></italic><italic>&#x03B8;</italic> represents the data used to compute the conditional PDF for row <italic>&#x03C6;</italic>, and <italic>f<sub>marginal</sub></italic> is the data used to compute the marginal distribution over the <italic>n<sub>&#x03B8;</sub></italic> rows of the Y axis.</p>
</sec>
</sec>
<sec id="c7-s4">
<label>7.4</label>
<title>Learned Sampling and PDF evaluation</title>
<p>We aim to find a learning-based representation useful to sample directions from an environment map following its distribution, as well as evaluating the probability density (PDF) for any given direction <italic>p</italic>(<italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>). While this could be approached using two independent neural networks, there is no guarantee that a sample generated by a network is drawn with the probability estimated by the other network, surely affecting the illumination integral and producing visual artifacts (<italic>e.g.</italic> fireflies). Neural networks by default are typically non-invertible and can become prohibitively expensive once model sizes grow, either by stacking more layers or adding more neurons. We instead leverage <italic>normalizing flows</italic>, defined by specifying a bijective function <inline-formula><mml:math id="M466" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> or its inverse <inline-formula><mml:math id="M467" display='block'><mml:msup><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. This enables both sampling and evaluation simultaneously and coherently with a single invertible neural network.</p>
<sec id="c7-s4.s1">
<label>7.4.1</label>
<title>Formalization</title>
<p>We define our normalizing flow as a differentiable transformation <inline-formula><mml:math id="M468" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> which maps from a random vector <inline-formula><mml:math id="M469" display='block'><mml:mi>z</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula>, sampled from a 2-dimensional uniform distribution <italic>p<sub>z</sub></italic>(<italic>z</italic>), to samples <inline-formula><mml:math id="M470" display='block'><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula>:</p>
<p><disp-formula id="Eq007-3"><label><target target-type="page" id="pges_195"/>(7.3)</label> <mml:math id="M471" display='block'><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="script">F</mml:mi><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo><mml:mtext>where</mml:mtext><mml:mi>z</mml:mi><mml:mo>&#x007E;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>z</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:math></disp-formula></p>
<p>The probability density of a direction <inline-formula><mml:math id="M472" display='block'><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="script">F</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="normal">&#x211D;</mml:mi></mml:math></inline-formula> under the normalizing flow <inline-formula><mml:math id="M473" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> is obtained through a change of variables. This can be computed analytically because <inline-formula><mml:math id="M474" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> is designed as an invertible function. Formally:</p>
<p><disp-formula id="Eq007-4"><label>(7.4)</label> <mml:math id="M475" display='block'><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="script">F</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>z</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="script">F</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>det</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03B4;</mml:mi><mml:msup><mml:mi mathvariant="script">F</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>&#x03B4;</mml:mi><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><inline-formula><mml:math id="M476" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> locally transforms the density of original noise distribution <italic>p<sub>z</sub></italic>(<italic>z</italic>). The intensity of this change if measured by the determinant of the Jacobian of the transformation, as in the right hand side of the equation above. <italic>p<sub>z</sub></italic>(<italic>z</italic>) is chosen to be a bidimensional uniform distribution bounded between 0 and <inline-formula><mml:math id="M477" display='block'><mml:mn>1</mml:mn><mml:mo>:</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>z</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="script">V</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:math></inline-formula>.</p>
<p>As shown, <inline-formula><mml:math id="M478" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> provides the two calculations required in MIS for Monte Carlo integration. However, <inline-formula><mml:math id="M479" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> needs to meet two competing requirements: it has to be as accurate as possible, <italic>i.e.</italic>, close to the real PDF, while being computationally efficient, both in evaluation and sampling. These competing requirements define our design choices, which we motivate next.</p>
</sec>
<sec id="c7-s4.s2">
<label>7.4.2</label>
<title>Design of <inline-formula><mml:math id="M480" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula></title>
<p>The design of <inline-formula><mml:math id="M481" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> follows three principles: First, to achieve accurate renderings, we need it to be expressive enough to represent any complex probability density distribution. Second, we need an analytic form to compute its inverse <inline-formula><mml:math id="M482" display='block'><mml:msup><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. Finally, for efficiency, computing <inline-formula><mml:math id="M483" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M484" display='block'><mml:msup><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and its Jacobian determinant needs to be as fast as possible.</p>
<p><italic>Coupling Flows</italic> [<xref ref-type="bibr" rid="CIT072">DKB14</xref>; <xref ref-type="bibr" rid="CIT073">DSB16</xref>; <xref ref-type="bibr" rid="CIT207">KD18</xref>] achieve the efficiency requirements by using a single feed-forward pass of <inline-formula><mml:math id="M485" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M486" display='block'><mml:msup><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and preserving a reduced level of computation, required to evaluate the Jacobian determinant, which is a lower triangular matrix for this type of flows. However, because all computation is done on a single pass to the models, coupling flows are typically less expressive than <italic>Autoregressive Flows</italic> [<xref ref-type="bibr" rid="CIT285">M&#x00FC;l+20</xref>; <xref ref-type="bibr" rid="CIT206">Kin+16</xref>; <xref ref-type="bibr" rid="CIT301">PPM17</xref>; <xref ref-type="bibr" rid="CIT167">Hua+18a</xref>; <xref ref-type="bibr" rid="CIT302">PSM19</xref>; <xref ref-type="bibr" rid="CIT204">Khe+21</xref>], which in turn, are less efficient during test time. We refer the reader to recent surveys [<xref ref-type="bibr" rid="CIT209">KPB20</xref>; <xref ref-type="bibr" rid="CIT300">Pap+21</xref>] for more comprehensive analyses of these topics.</p>
<p>To balance between efficiency and accuracy, we design <inline-formula><mml:math id="M487" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> as a <italic>Coupling Flow</italic> and incorporate expressiveness through the coupling layer design. To this end, we evaluate several design choices available in the literature in the context of our problem. Recent work include <italic>Affine Coupling</italic> layers [<xref ref-type="bibr" rid="CIT073">DSB16</xref>], <italic>Piecewise Linear</italic> and <italic>Piecewise Quadratic</italic> [<xref ref-type="bibr" rid="CIT284">M&#x00FC;l+19</xref>], <italic>Cubic-Spline Flows</italic> [<xref ref-type="bibr" rid="CIT081">Dur+19b</xref>] and <italic>Rational-Quadratic Spline Flows</italic> [<xref ref-type="bibr" rid="CIT080">Dur+19a</xref>]. The type of coupling controls the <target target-type="page" id="pges_196"/>complexity of <inline-formula><mml:math id="M488" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>, which can further be tuned by modifying the number of <italic>coupling layers</italic>, the <italic>width</italic> and <italic>depth</italic> of the neural network that define each coupling, and the number of <italic>bins</italic> which define the complexity of the polynomial function inside each coupling layer.</p>
<p>In addition, as in any neural network, it is possible to tweak the internal normalization layers, activation functions, optimizer type and configuration, regularization choices, etc. These create a combinatorial explosion of design choices, with two competing objectives of accuracy and efficiency. We found that <italic>Rational-Quadratic Spline Flows</italic> [<xref ref-type="bibr" rid="CIT080">Dur+19a</xref>] with a reduced amount of small coupling layers are expressive to a sufficient degree for every environment map that we tested, while being very efficient during evaluation, achieving orders of magnitude less computational cost than tabulated sampling. Because <inline-formula><mml:math id="M489" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> is efficient and fully differentiable, it could potentially be incorporated into inverse rendering problems. We specify our model design choices in <xref ref-type="sec" rid="c7-s6">Section 7.6</xref> and evaluate them in <xref ref-type="sec" rid="c7-s8.s1">Section 7.8.1</xref>.</p>
</sec>
<sec id="c7-s4.s3">
<label>7.4.3</label>
<title>Training <inline-formula><mml:math id="M490" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula></title>
<p>We train our normalizing flow to minimize the aggregated negative log-likelihood of a set of samples, <inline-formula><mml:math id="M491" display='block'><mml:mi mathvariant="script">B</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup></mml:math></inline-formula>, that are drawn dynamically in batches of size <italic>N</italic> from the input</p>
<p>environment map <inline-formula><mml:math id="M492" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula> as described in <xref ref-type="sec" rid="c7-s3.s2">Section 7.3.2</xref>.</p>
<p>During each training step, we minimize the average negative log likelihood across the elements in the batch <inline-formula><mml:math id="M493" display='block'><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mi>b</mml:mi><mml:mi>N</mml:mi></mml:munderover><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>nll</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> where:</p>
<p><disp-formula id="Eq007-5"><label>(7.5)</label> <mml:math id="M494" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="script">F</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>z</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="script">F</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>det</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03B4;</mml:mi><mml:msup><mml:mi mathvariant="script">F</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>&#x03B4;</mml:mi><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This is done by learning to map from the distribution of samples to the noise distribution <italic>p<sub>z</sub></italic>(<italic>z</italic>). Given a sufficient number of training batches, <inline-formula><mml:math id="M495" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> learns to accurately sample from the distribution of samples encoded in the input environment map <inline-formula><mml:math id="M496" display='block'><mml:mi mathvariant="script">I</mml:mi></mml:math></inline-formula>. Because <inline-formula><mml:math id="M497" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> is invertible and due to our efficient model design, the probability of a sample (<italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>) can easily be computed by its inverse <inline-formula><mml:math id="M498" display='block'><mml:msup><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<p>During training, we observe that most of the computation time is used to generate the ground truth batches of samples <inline-formula><mml:math id="M499" display='block'><mml:mi mathvariant="script">B</mml:mi></mml:math></inline-formula>, which is the very process that we are interested in improving with <inline-formula><mml:math id="M500" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>. We also notice that achieving a training procedure that works for every environment map is challenging due to gradient instabilities. We introduce a series of modifications to the training procedure to allow for stable training dynamics and avoid NANs for any input environment map, which we specify in <xref ref-type="sec" rid="c7-s6">Section 7.6</xref>.</p>
</sec>
</sec>
<sec id="c7-s5">
<label><target target-type="page" id="pges_197"/>7.5</label>
<title>Environment Map Compression</title>
<p>In addition to learning to sample and evaluate the PDF, we train a compression neural network <inline-formula><mml:math id="M501" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> which learns to map from directions to RGB values: <inline-formula><mml:math id="M502" display='block'><mml:mi mathvariant="script">C</mml:mi><mml:mo>(</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mi>G</mml:mi><mml:mi>B</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo>&#x00A0;</mml:mo><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mi>G</mml:mi><mml:mi>B</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211D;</mml:mi><mml:mrow><mml:mo>+</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. This serves two purposes: First, it provides a memory and time efficient representation of the environment map, thus further improving the render times and computational cost. Second, it transforms the tabulated representation into a continuous and differentiable function for which gradients can be computed, which may prove useful for inverse rendering scenarios. For the model design of <inline-formula><mml:math id="M503" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula>, we build upon previous work on implicit neural representations. In particular, neural network architectures that work on the frequency domain have shown increased performance with respect to MLP for modeling natural signals [<xref ref-type="bibr" rid="CIT384">Tan+20</xref>; <xref ref-type="bibr" rid="CIT366">Sit+20</xref>], both in reconstruction quality and parameter efficiency. We thus model <inline-formula><mml:math id="M504" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> as a shallow <italic>sinusoidal MLP</italic>, a <italic>SIREN</italic> [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>]. This architectural design has been explored in recent work on environment map approximation, such as <italic>RENI</italic> [<xref ref-type="bibr" rid="CIT105">GES22</xref>]. However, as we do not aim to learn any prior over natural illumination, we can remove some constraints introduced in <italic>RENI</italic>, most notably the <italic>rotation equivariance</italic> requirement. Further, we do not use any latent vector to condition the output of our model, nor we do require <italic>Vector Neurons</italic> for our representation. With our simplifications, we achieve much higher reconstruction quality with the same parameter count. We train <inline-formula><mml:math id="M505" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> on linear RGB space, with a pixel-wise reconstruction loss. Specifically, we use an <inline-formula><mml:math id="M506" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> loss, as it produces sharper and more accurate reconstructions than higher-order norms, such as <inline-formula><mml:math id="M507" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> [<xref ref-type="bibr" rid="CIT176">Iso+17</xref>; <xref ref-type="bibr" rid="CIT338">RG21</xref>]. We also observe strong gradient instabilities when training with <inline-formula><mml:math id="M508" display='block'><mml:msub><mml:mrow><mml:mi mathvariant="script">L</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> on linear RGB. As <inline-formula><mml:math id="M509" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> expects inputs in radians, but the input image has ranges of <italic>H</italic> &#x00D7; <italic>W</italic> pixels, our full loss becomes:</p>
<p><disp-formula id="Eq007-6"><label>(7.6)</label> <mml:math id="M510" display='block'><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mtext>recon</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mi>i</mml:mi><mml:mi>H</mml:mi></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mi>j</mml:mi><mml:mi>W</mml:mi></mml:munderover><mml:msub><mml:mrow><mml:mo>&#x2225;</mml:mo><mml:mi mathvariant="script">C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>&#x03C0;</mml:mi></mml:mrow><mml:mi>H</mml:mi></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi></mml:mrow><mml:mi>W</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi mathvariant="script">I</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo><mml:mo>&#x2225;</mml:mo></mml:mrow><mml:mn>1</mml:mn></mml:msub></mml:math></disp-formula></p>
<p>We use mini-batching to train our compression models <inline-formula><mml:math id="M511" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula>, using uniformly sampled directions <italic>&#x03B8;</italic>, <italic>&#x03C6;</italic>. Previous work [<xref ref-type="bibr" rid="CIT105">GES22</xref>; <xref ref-type="bibr" rid="CIT380">Szt+21</xref>; <xref ref-type="bibr" rid="CIT337">Rod+23a</xref>; <xref ref-type="bibr" rid="CIT229">Lav+21</xref>] introduce adaptive sampling, cosine weighting or specular peak attenuation to similar loss functions for environment map or BRDF reconstruction. However, for our particular problem, we empirically observed that these additions tend to contribute negatively either to the training dynamics, or to the final quality of the reconstruction.</p>
</sec>
<sec id="c7-s6">
<label>7.6</label>
<title>Implementation Details</title>
<p><bold>Preprocessing</bold> Before the rotation algorithm described In <xref ref-type="sec" rid="c7-s3.s1">Section 7.3.1</xref>, we resize every input HDRi environment map to a resolution of (2000, 4000) pixels using area interpolation.</p>
<p><target target-type="page" id="pges_198"/>This is an optional step, but it helps us standardize our experiments so that input resolution does not change training times. Note that evaluation times of our NEnv models depend on the model sizes and parameterizations, not on the resolution they were trained on.</p>
<p><bold>Model Design and Training</bold> Our models sizes, loss functions, coupling layer design, normalization, optimizer and training configurations were selected using a combination of manual tuning and Bayesian hyperparameter optimization using <italic>Weights and Biases</italic> [<xref ref-type="bibr" rid="CIT032">Bie20</xref>] for a variety of representative environment maps.</p>
<p>We train the models using PyTorch [<xref ref-type="bibr" rid="CIT304">Pas+19</xref>] 1.11 and Torchvision [<xref ref-type="bibr" rid="CIT262">MR10</xref>]. Our flow and compression networks are trained independently, and individually for each input environment map. We empirically observe that mixed precision training [<xref ref-type="bibr" rid="CIT274">Mic+18</xref>] introduces strong training instabilities for this problem, so we train the models using <italic>float32</italic> precision. However, our models can be evaluated using half precision, which significantly increases their efficiency during test time. <inline-formula><mml:math id="M512" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M513" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> each use around 3 MBs each, while a (2000, 4000) resolution HDRi environment map uses around 25 MBs.</p>
<p><bold>Normalizing Flow</bold> Our normalizing flows <inline-formula><mml:math id="M514" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> build upon the official implementation in [<xref ref-type="bibr" rid="CIT080">Dur+19a</xref>]. We use Batch Normalization [<xref ref-type="bibr" rid="CIT175">IS15</xref>] in the residual layers of our flows. To maximize evaluation efficiency, our flows have a limited number of trainable parameters. In particular, we use only 2 coupling layers, with 2 hidden layers, each with 256 hidden units and 256 <italic>bins</italic>. We use spline flows as our coupling layer type [<xref ref-type="bibr" rid="CIT080">Dur+19a</xref>]. Our initial learning rate is of 5<italic>e</italic> &#x2212; 4, which is halved every 2500 iterations. We use Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>] as our optimizer. To stabilize training, we use gradient norm clipping [<xref ref-type="bibr" rid="CIT445">Zha+20</xref>] (<italic>maxnorm</italic> = 1). This is crucial for well-behaved models and efficient convergence. We train our models with a batch size of 100000 for 15000 epochs, which takes around 2 hours on an Nvidia RTX 3060 GPU. Note that fewer iterations (eg 5000) are typically sufficient for adequate results, while longer training helps to resolve additional details.</p>
<p><bold>Compression Network</bold> We use a small sinusoidal MLP as our environment map compression network <inline-formula><mml:math id="M515" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula>. We use the official implementation in [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>]. We train our models with a batch size of 500000 of randomly sampled directions and linear RGB values, for 10000 epochs, which takes around 0.5 hours on an Nvidia RTX 3060 GPU. Our initial learning rate is 5<italic>e</italic> &#x2212; 4, which is halved every 2000 iterations. We use Adam [<xref ref-type="bibr" rid="CIT205">KB15</xref>] as our optimizer. To stabilize training, we use gradient norm clipping (<italic>maxnorm</italic> = 50). Our compression networks have 3 layers with 512 hidden units each. We initialize their weights as described in [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>].</p>
<p><bold>Rendering</bold> We use our own path tracer (PT) rendering pipeline as a baseline engine. This PT uses an implementation inspired by <italic>Wavefront</italic> [<xref ref-type="bibr" rid="CIT397">Van11</xref>; <target target-type="page" id="pges_199"/><xref ref-type="bibr" rid="CIT226">LKA13</xref>], powered by <italic>Embree</italic> [<xref ref-type="bibr" rid="CIT404">Wal+14</xref>]. We distribute render work in tiles of (25,25) pixels that perform path&#x015B;tracing in usual Wave-front steps (ray generation, intersection, shading, connection), each tile using a separate CPU thread. Note that tile size could be changed, as our selected tile dimensions are fixed due to hardware limitations and the amount of current thread active simultaneously (Intel Core i7-7700K CPU @ 4.20GHz, 8 threads).</p>
<p>Under this architecture, we perform our environment map sampling step once per tile, before every pixel in each tile performs its connection step. This way, we can either use a CPU approach [<xref ref-type="bibr" rid="CIT053">Cla+21</xref>] per pixel or integrate a single call to <inline-formula><mml:math id="M516" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> to obtain all samples that are required in the current tile. We use <italic>LibTorch</italic> to use our PyTorch implementation from our C++ code. Intersections are performed in a single call using a ray packet call from Embree. Integrating <inline-formula><mml:math id="M517" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> is more challenging, since we require ray data. However, we can store all this information simultaneously: We prepare all rays for intersection, and perform a single call to <inline-formula><mml:math id="M518" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula>. This way, we obtain all values at once. Note that these evaluations will be discarded if the next path&#x015B;tracing bounce does not go out of bounds.</p>
</sec>
<sec id="c7-s7">
<label>7.7</label>
<title>Dataset</title>
<p>Our goal is to propose a method that works for any type of global illumination which can be represented using environment maps. To this end, we gather a dataset of 32 high resolution HDRi environment maps from publicly available sources, notably HDRMaps and Poly Haven, to thoroughly evaluate our models. Our dataset contains natural daylight, including sunny midday illumination, cloudy diffuse images, sunset, and dusk. Further, we also test our method on indoor scenes, from studio lighting, to a concert hall or cathedral illuminations. We purposefully include very challenging cases with multiple colored light sources, which helps us understand the limitations and capabilities of previous work and our models.</p>
</sec>
<sec id="c7-s8">
<label>7.8</label>
<title>Evaluation</title>
<p>In this section, we evaluate our models, both quantitatively and qualitatively, We first measure their performance in terms of reconstructing the probability density and the RGB values of the input environment maps, and in computational cost. Finally, we measure their integrated quality in final rendered images compared to a variety of baselines. Unless stated otherwise, we use the full dataset described in <xref ref-type="sec" rid="c7-s7">Section 7.7</xref> for our analyses.</p>
<p><target target-type="page" id="pges_200"/>To make comparisons fair, we use the following configuration: For RENI [<xref ref-type="bibr" rid="CIT105">GES22</xref>], we train a new model from scratch with the exact same architecture design as our <inline-formula><mml:math id="M520" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula>, and train it for 3000 epochs. For Spherical Harmonics (SH) [<xref ref-type="bibr" rid="CIT356">See66</xref>; <xref ref-type="bibr" rid="CIT320">RH01b</xref>], we use 1024 coefficients, and for Spherical Gaussians (SG) [<xref ref-type="bibr" rid="CIT429">Xu+13</xref>], 16 &#x00D7; 32 dimensions. These configurations have a larger amount of parameters than what is typically used for these methods, especially in real time rendering. However, our goal is to maximize the reconstruction quality for each method and environment map, even if it results in an increased memory footprint. For the three methods we evaluate, the output resolution is set to (256, 512) pixels. We follow the official implementation of [<xref ref-type="bibr" rid="CIT105">GES22</xref>] for all these comparisons.</p>
<sec id="c7-s8.s1">
<label>7.8.1</label>
<title>PDF Fit Accuracy</title>
<p>To validate our design choices for the coupling layers of our normalizing flow <inline-formula><mml:math id="M521" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>, we train different versions of <inline-formula><mml:math id="M522" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> for every environment map in our dataset, exclusively changing the type of polynomial function in the coupling layer. We measure the <inline-formula><mml:math id="M523" display='block'><mml:mi mathvariant="script">X</mml:mi></mml:math></inline-formula> <sup>2</sup> distance, as well as the KL divergence <inline-formula><mml:math id="M524" display='block'><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="script">F</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi mathvariant="script">I</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="CIT215">KL51</xref>] between the probabilities encoded by our normalizing flow <inline-formula><mml:math id="M525" display='block'><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="script">F</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and the ground truth counterparts <inline-formula><mml:math id="M526" display='block'><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="script">I</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> as follows:</p>
<p><disp-formula id="Eq007-7"><label>(7.7)</label> <mml:math id="M527" display='block'><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="script">F</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi mathvariant="script">I</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>&#x03C0;</mml:mi></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="script">F</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="script">F</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="script">I</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This metric, also used in [<xref ref-type="bibr" rid="CIT284">M&#x00FC;l+19</xref>], measures the distance between the learned probability distribution and the real PDF encoded in the environment map. In <xref ref-type="fig" rid="fig-7.5">Figure 7.5</xref>, we show the aggregated divergences for our dataset for three different coupling layer types. As shown, the <italic>Piecewise Linear</italic> coupling proposed in [<xref ref-type="bibr" rid="CIT284">M&#x00FC;l+19</xref>] is outperformed by its more expressive <italic>Piecewise Quadratic</italic> variation, while the <italic>Rational-Quadratic Spline Flows</italic> proposed in [<xref ref-type="bibr" rid="CIT080">Dur+19a</xref>] achieve the lowest divergence overall, albeit by a relatively small margin. This improvement is achieved solely by adding expressiveness to the coupling layers, without any additional parameter cost. We provide a qualitative evaluation in <xref ref-type="fig" rid="fig-7.6">Figure 7.6</xref>, where we show that linear and quadratic coupling layers encode less detailed probability distributions. These differences are particularly relevant for inputs with multiple distant light sources.</p>
<fig id="fig-7.5">
<label>Figure 7.5:</label>
<caption><title>PDF <inline-formula><mml:math id="M519" display='block'><mml:mi mathvariant="script">X</mml:mi></mml:math></inline-formula> <sup>2</sup> and KL Divergence [<xref ref-type="bibr" rid="CIT215">KL51</xref>] of normalizing flows using different coupling layer types, including Linear and Quadratic Couplings from [<xref ref-type="bibr" rid="CIT284">M&#x00FC;l+19</xref>] and Spline Layers [<xref ref-type="bibr" rid="CIT080">Dur+19a</xref>].</title></caption>
<alt-text>PDF X 2 and KL Divergence [KL51] of normalizing flows using different coupling layer types, including Linear and Quadratic Couplings from [M&#x00FC;l+19] and Spline Layers [Dur+19a].</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.5.jpg"/>
</fig>
<fig id="fig-7.6">
<label><target target-type="page" id="pges_201"/>Figure 7.6:</label>
<caption><title>A qualitative comparison between the type of coupling layer used in our normalizing flows. On the top two rows, we show the input environment maps and their ground truth PDFs. Using the linear and quadratic couplings defined in [<xref ref-type="bibr" rid="CIT284">M&#x00FC;l+19</xref>] we achieve somewhat accurate encodings. With spline flows [<xref ref-type="bibr" rid="CIT080">Dur+19a</xref>], we achieve sharper and more accurate probability distributions. We use a &#x03B3; = 2.2 to tonemap the RGB and PDF maps to help visualization. We highlight relevant regions using blue insets. Best viewed in color on a screen.</title></caption>
<alt-text>A qualitative comparison between the type of coupling layer used in our normalizing flows. On the top two rows, we show the input environment maps and their ground truth PDFs. Using the linear and quadratic couplings defined in [M&#x00FC;l+19] we achieve somewhat accurate encodings. With spline flows [Dur+19a], we achieve sharper and more accurate probability distributions. We use a &#x03B3; = 2.2 to tonemap the RGB and PDF maps to help visualization. We highlight relevant regions using blue insets. Best viewed in color on a screen.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.6.jpg"/>
</fig>
<p>We also evaluate whether a larger model can yield significant improvements in accuracy. In <xref ref-type="fig" rid="fig-7.7">Figure 7.7</xref>, we show the results of a large flow (<inline-formula><mml:math id="M528" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>-XL), with twice as many layers and neurons per layer, a total of approximately 7 times more trainable parameters and model size in memory. Even though it indeed reduces the divergence from <inline-formula><mml:math id="M529" display='block'><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="script">F</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi mathvariant="script">I</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula> = 0.308 to <inline-formula><mml:math id="M530" display='block'><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="script">F</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>XL</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi mathvariant="script">I</mml:mi><mml:mo>)</mml:mo></mml:math></inline-formula> = 0.281, this model takes approximately one day to train and makes sampling 3.43 times slower. This larger model only <target target-type="page" id="pges_202"/>achieves a small gain in render accuracy (from 0.059 to 0.048 in FLIP [<xref ref-type="bibr" rid="CIT011">And+20</xref>]), suggesting that larger models exhibit diminishing returns in terms of quality. We believe our chosen model sizes adequately balance accuracy and speed. However, for practitioners that would like better accuracy at the cost of computational performance, we suggest using more coupling layers and more bins, as they tend to be the most effective design decisions beyond the coupling type.</p>
<fig id="fig-7.7">
<label>Figure 7.7:</label>
<caption><title>A visualization of the impact of the size of <inline-formula><mml:math id="M531" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>. <italic>NEnv-XL</italic> is a larger version of <inline-formula><mml:math id="M532" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>, with twice as many layers, bins and neurons per layer. This yields on small improvements on render and reconstruction quality, at the cost of sampling efficiency.</title></caption>
<alt-text>A visualization of the impact of the size of F. NEnv-XL is a larger version of F, with twice as many layers, bins and neurons per layer. This yields on small improvements on render and reconstruction quality, at the cost of sampling efficiency.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.7.jpg"/>
</fig>
</sec>
<sec id="c7-s8.s2">
<label>7.8.2</label>
<title>Environment Map Reconstruction</title>
<p>We now aim to measure the accuracy of the environment maps encoded by <inline-formula><mml:math id="M533" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> and previous work on this problem. In <xref ref-type="table" rid="c7-tab1">Table 7.1</xref>, we show the average error on a variety of metrics for every environment map in our dataset. We use pixel-wise (PSNR) and perceptual metrics (SSIM [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>], LPIPS-VGG [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>], FLIP [<xref ref-type="bibr" rid="CIT011">And+20</xref>]) to thoroughly understand the differences between each method. We compute these metrics using <italic>PIQ</italic> [<xref ref-type="bibr" rid="CIT199">KZP19</xref>] and the implementation in [<xref ref-type="bibr" rid="CIT011">And+20</xref>], on images of (512, 256) pixels. We apply a &#x03B3; = 2.2 to tonemap them before computing the distances. As shown, our network <inline-formula><mml:math id="M534" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> clearly outperforms every other method, with SH struggling significantly, while RENI and SG achieve comparably better performances. We also provide a qualitative comparison in <xref ref-type="fig" rid="fig-7.8">Figure 7.8</xref>, for a subset of our dataset, where it can be seen that <inline-formula><mml:math id="M535" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> provides high-quality <target target-type="page" id="pges_203"/>reconstructions. RENI is optimized for outdoor, natural lighting, and it struggles with indoor illumination, providing overly smooth outputs. SG provides sharper results, while SH yields visually unpleasant results for every case: it is well known that the cyclic nature of SH tends to produce ringing artifacts, especially for high-contrast illumination environments and a large number of coefficients. We study the impact of these reconstructions on final renders in <xref ref-type="sec" rid="c7-s8.s4">Section 7.8.4</xref>. Our representation <inline-formula><mml:math id="M537" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> has the additional advantage that, thanks to its reconstruction quality, it can be used directly as background for samples that do not hit any geometry on the scene, while other alternatives require additional storage for higher quality versions, or even the original image.</p>
<table-wrap id="c7-tab1">
<label>Table 7.1:</label>
<caption><title>Average (&#x00B1; std.) environment map reconstruction error for different methods, with pixel-wise and perceptual metrics. We use a color code to highlight <bold>best</bold> and <bold>worst</bold> cases.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left"><p>Method</p></th>
<th valign="top" align="left"><p>SH [<xref ref-type="bibr" rid="CIT320">RH01b</xref>]</p></th>
<th valign="top" align="left"><p>SG [<xref ref-type="bibr" rid="CIT429">Xu+13</xref>]</p></th>
<th valign="top" align="left"><p>RENI [<xref ref-type="bibr" rid="CIT105">GES22</xref>]</p></th>
<th valign="top" align="left"><p><bold>NEnv</bold> <inline-formula><mml:math id="M536" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> <bold>(Ours)</bold></p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p>PSNR &#x2191;</p></td>
<td valign="top" align="left"><p><bold>13.68</bold>&#x00B1;4.88</p></td>
<td valign="top" align="left"><p>20.54&#x00B1;5.04</p></td>
<td valign="top" align="left"><p>18.92&#x00B1;3.24</p></td>
<td valign="top" align="left"><p><bold>30.65</bold>&#x00B1;5.59</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SSIM [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>] &#x2191;</p></td>
<td valign="top" align="left"><p><bold>0.276</bold>&#x00B1;0.20</p></td>
<td valign="top" align="left"><p>0.559&#x00B1;0.16</p></td>
<td valign="top" align="left"><p>0.557&#x00B1;0.16</p></td>
<td valign="top" align="left"><p><bold>0.873</bold>&#x00B1;0.08</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>LPIPS [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] &#x2193;</p></td>
<td valign="top" align="left"><p><bold>0.687</bold>&#x00B1;0.07</p></td>
<td valign="top" align="left"><p>0.637&#x00B1;0.12</p></td>
<td valign="top" align="left"><p>0.657&#x00B1;0.12</p></td>
<td valign="top" align="left"><p><bold>0.155</bold>&#x00B1;0.08</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>FLIP [<xref ref-type="bibr" rid="CIT011">And+20</xref>] &#x2193;</p></td>
<td valign="top" align="left"><p><bold>0.505</bold>&#x00B1;0.21</p></td>
<td valign="top" align="left"><p>0.248&#x00B1;0.11</p></td>
<td valign="top" align="left"><p>0.278&#x00B1;0.09</p></td>
<td valign="top" align="left"><p><bold>0.062</bold>&#x00B1;0.03</p></td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-7.8">
<label>Figure 7.8:</label>
<caption><title>Environment maps encoded by our method and previous work, including RENI [<xref ref-type="bibr" rid="CIT105">GES22</xref>], Spherical Gaussian (SG) [<xref ref-type="bibr" rid="CIT429">Xu+13</xref>] and Spherical Harmonics (SH) [<xref ref-type="bibr" rid="CIT320">RH01b</xref>]. We use a &#x03B3; = 2.2 to tonemap the RGB maps to help visualization.</title></caption>
<alt-text>Environment maps encoded by our method and previous work, including RENI [GES22], Spherical Gaussian (SG) [Xu+13] and Spherical Harmonics (SH) [RH01b]. We use a &#x03B3; = 2.2 to tonemap the RGB maps to help visualization.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.8.jpg"/>
</fig>
</sec>
<sec id="c7-s8.s3">
<label><target target-type="page" id="pges_204"/>7.8.3</label>
<title>Computational Cost</title>
<p>In <xref ref-type="fig" rid="fig-7.9">Figure 7.9</xref>, we show the average time per sample obtained by the previous work and our normalizing flow <inline-formula><mml:math id="M538" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>. On average, <italic>Hierarchical Warping</italic> [<xref ref-type="bibr" rid="CIT053">Cla+21</xref>; <xref ref-type="bibr" rid="CIT306">Pha19</xref>] obtains <inline-formula><mml:math id="M539" display='block'><mml:mn>6.48</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00B1;</mml:mo><mml:mn>1.84</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> milliseconds (ms) per sample, with the <italic>Piecewise&#x015B;Constant</italic> method [<xref ref-type="bibr" rid="CIT308">PJH16</xref>] achieving <inline-formula><mml:math id="M540" display='block'><mml:mn>1.12</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00B1;</mml:mo><mml:mn>3.24</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> ms, while <inline-formula><mml:math id="M541" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> obtains <inline-formula><mml:math id="M542" display='block'><mml:mn>6.83</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00B1;</mml:mo><mml:mn>3.69</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. To make these comparisons more favorable to each method, the timings for <italic>Hierarchical Warping</italic> [<xref ref-type="bibr" rid="CIT053">Cla+21</xref>; <xref ref-type="bibr" rid="CIT306">Pha19</xref>] and <italic>Piecewise&#x015B; Constant</italic> [<xref ref-type="bibr" rid="CIT307">PJH20</xref>] are measured on CPU, since both algorithms rely on binary search which is difficult to parallelize on GPU, while the timings for NEnv are measured on GPU. For <inline-formula><mml:math id="M543" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>, we use a sampling batch size of 200000, and we also include the times required to upload and download the samples from GPU, as in a real rendering scenario. We measure these times on the same hardware to make results comparable.</p>
<fig id="fig-7.9">
<label>Figure 7.9:</label>
<caption><title>On the left, average seconds per sample for the baseline sampling algorithms and our model <inline-formula><mml:math id="M544" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>. On the right, evaluation times for a tabulated baseline and our <inline-formula><mml:math id="M545" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula>. Note that we use a logarithmic scale for this plot.</title></caption>
<alt-text>On the left, average seconds per sample for the baseline sampling algorithms and our model F. On the right, evaluation times for a tabulated baseline and our c. Note that we use a logarithmic scale for this plot.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.9.jpg"/>
</fig>
<p>While being differentiable and losing little reconstruction accuracy, <inline-formula><mml:math id="M546" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> achieves, on average,</p>
<p>16.83 times more samples per second than the <italic>Piecewise-Constant</italic> method and 94.85 than <italic>Hierarchical Warping</italic>, with a maximum difference of 106.06. Furthermore, the timings for the <italic>Piecewise-Constant</italic> and <italic>Hierarchical Warping</italic> methods vary up to 15% and 12.5% respectively depending on the content of the environment map, while the variance of <inline-formula><mml:math id="M547" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> is negligible. Not only did we achieve up to two orders of magnitude faster sampling than previous work, our method was also significantly more consistent in terms of timings. We found that <italic>Hierarchical Warping</italic> and the <italic>Piecewise-Constant</italic> methods typically struggle more with environment maps with multiple, distant light sources, while they are more efficient for natural daylight illumination. <target target-type="page" id="pges_205"/>In terms of RGB evaluation, a tabulated version of the environment map requires an average of <inline-formula><mml:math id="M548" display='block'><mml:mn>8.43</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00B1;</mml:mo><mml:mn>1.44</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> ms per direction, while our model <inline-formula><mml:math id="M549" display='block'><mml:mn>2.83</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00B1;</mml:mo><mml:mn>7.85</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> ms. Note that we use <italic>TorchScript</italic> with <italic>Pytorch 1.11</italic> to measure the times for <inline-formula><mml:math id="M550" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M551" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula>. Newer versions of these libraries, which allow for compiled models, will likely help achieve faster evaluations without changing the model architectures.</p>
</sec>
<sec id="c7-s8.s4">
<label>7.8.4</label>
<title>Rendering Comparisons</title>
<p>The goal of NEnv is to generate accurate yet fast environment map representations that do not make compromises in terms of rendering quality. We use the <italic>Piecewise-Constant</italic> PDFs as our ground truth (GT) and perform qualitative and quantitative analyses to show the reconstruction quality of NEnv with respect to previous work. We provide a fine grained analysis of NEnv, in which we evaluate each of the models separately and a full version of the model which uses both <inline-formula><mml:math id="M552" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M553" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula>. We use our path tracing engine configured as explained in <xref ref-type="sec" rid="c7-s6">Section 7.6</xref> to render the same scene for each environment map. Each scene is rendered using a simple pinhole camera (no Depth of Field) and the environment maps as the only source of light, so we can reduce as much as possible external sources of Monte Carlo noise at low sampling. We use MIS, so both sampling and PDF functions for our maps impact the resulting image. All images are rendered with a resolution of (1080, 1920) pixels, 32 samples, maximum depth of 16, and path throughput weight Russian Roulette [<xref ref-type="bibr" rid="CIT308">PJH16</xref>].</p>
<p>In <xref ref-type="fig" rid="fig-7.10">Figure 7.10</xref>, we show a qualitative comparison between different methods on the same scene, illuminated with different environment maps. We use an object from [<xref ref-type="bibr" rid="CIT355">Sch+17</xref>], which helps in highlighting the differences between far field illumination representations. As shown, our method, in any of their configurations, generates images that closely match the ground truth, while previous work struggles significantly, particularly in indoor illumination. Finally, in <xref ref-type="table" rid="c7-tab2">Table 7.2</xref>, we show a quantitative comparison of the average error obtained by these methods, for our entire dataset. For this analysis, we only measure the error on the object and floor on the renders in <xref ref-type="fig" rid="fig-7.10">Figure 7.10</xref>, masking out the background. We observe that, even when combining our sampling and compression networks, our reconstruction quality is much higher than any of the methods in the previous work. Interestingly, the differences in environment map reconstruction quality we measured in <xref ref-type="table" rid="c7-tab1">Table 7.1</xref>, which were very unfavorable to SH, are not exactly correlated to rendering error, suggesting that SH, while being a pixel-wise inaccurate image representation, it tends to be precise on the more relevant information of the image, <italic>i.e.</italic>, where the key light sources lie. In terms of rendering, neither of the previous methods that we tested behaved consistently better than any other across every metric.</p>
<fig id="fig-7.10">
<label>Figure 7.10:</label>
<caption><title>Render comparisons of different configurations of our work with the ground truth environment map and previous work. Best viewed in color on a screen. Please zoom in for details.</title></caption>
<alt-text>Render comparisons of different configurations of our work with the ground truth environment map and previous work. Best viewed in color on a screen. Please zoom in for details.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-7.10.jpg"/>
</fig>
<table-wrap id="c7-tab2">
<label>Table 7.2:</label>
<caption><title>Average (&#x00B1; std.) render reconstruction error for different methods, with pixel-wise and perceptual metrics. We use a color code to highlight <bold>best</bold> and <bold>worst</bold> cases.</title></caption>
<table frame="box" rules="all">
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="left"><p>PSNR &#x2191;</p></th>
<th valign="top" align="left"><p>SSIM [<xref ref-type="bibr" rid="CIT419">Wan+04</xref>] &#x2191;</p></th>
<th valign="top" align="left"><p>LPIPS [<xref ref-type="bibr" rid="CIT450">Zha+18</xref>] &#x2193;</p></th>
<th valign="top" align="left"><p>FLIP [<xref ref-type="bibr" rid="CIT011">And+20</xref>] &#x2193;</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p>SH [<xref ref-type="bibr" rid="CIT320">RH01b</xref>]</p></td>
<td valign="top" align="left"><p>21.77&#x00B1;5.35</p></td>
<td valign="top" align="left"><p><bold>0.933</bold>&#x00B1;0.05</p></td>
<td valign="top" align="left"><p><bold>0.160</bold>&#x00B1;0.03</p></td>
<td valign="top" align="left"><p>0.150&#x00B1;0.06</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SG [<xref ref-type="bibr" rid="CIT429">Xu+13</xref>]</p></td>
<td valign="top" align="left"><p>21.62&#x00B1;6.34</p></td>
<td valign="top" align="left"><p>0.940&#x00B1;0.05</p></td>
<td valign="top" align="left"><p>0.153&#x00B1;0.03</p></td>
<td valign="top" align="left"><p><bold>0.154</bold>&#x00B1;0.07</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>RENI [<xref ref-type="bibr" rid="CIT105">GES22</xref>]</p></td>
<td valign="top" align="left"><p><bold>21.53</bold>&#x00B1;5.62</p></td>
<td valign="top" align="left"><p>0.936&#x00B1;0.04</p></td>
<td valign="top" align="left"><p>0.159&#x00B1;0.03</p></td>
<td valign="top" align="left"><p>0.151&#x00B1;0.05</p></td>
</tr>
<tr>
<td valign="top" align="left"><p><bold>NEnv (</bold><inline-formula><mml:math id="M554" display='block'><mml:mi mathvariant="script">F</mml:mi></mml:math></inline-formula> <bold>)</bold></p></td>
<td valign="top" align="left"><p>37.97&#x00B1;3.73</p></td>
<td valign="top" align="left"><p>0.996&#x00B1;0.00</p></td>
<td valign="top" align="left"><p>0.024&#x00B1;0.01</p></td>
<td valign="top" align="left"><p><bold>0.032</bold>&#x00B1;0.02</p></td>
</tr>
<tr>
<td valign="top" align="left"><p><bold>NEnv (</bold><inline-formula><mml:math id="M555" display='block'><mml:mi>c</mml:mi></mml:math></inline-formula><bold>)</bold></p></td>
<td valign="top" align="left"><p><bold>41.05</bold>&#x00B1;8.89</p></td>
<td valign="top" align="left"><p><bold>0.997</bold>&#x00B1;0.01</p></td>
<td valign="top" align="left"><p><bold>0.016</bold>&#x00B1;0.01</p></td>
<td valign="top" align="left"><p>0.036&#x00B1;0.02</p></td>
</tr>
<tr>
<td valign="top" align="left"><p><bold>NEnv (Both)</bold></p></td>
<td valign="top" align="left"><p>35.16&#x00B1;4.97</p></td>
<td valign="top" align="left"><p>0.994&#x00B1;0.01</p></td>
<td valign="top" align="left"><p>0.033&#x00B1;0.01</p></td>
<td valign="top" align="left"><p>0.049&#x00B1;0.02</p></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="c7-s9">
<label><target target-type="page" id="pges_206"/>7.9</label>
<title>Conclusions</title>
<p>In this work, we have presented a neural rendering method for joint compression, evaluation, and sampling of environment maps for global illumination. We have validated our models quantitatively using a diverse dataset of environment maps. Our proposed lightweight normalizing flows provide accurate yet efficient sampling and pdf evaluation, achieving significantly faster computational times than analytical approaches with negligible loss in quality. Further, our sinusoidal compression networks achieve high-quality environment map approximations, surpassing previous work in both quality and generality. Every component in our models is differentiable, and they can be seamlessly integrated into engines like Mitsuba [<xref ref-type="bibr" rid="CIT297">Nim+19</xref>; <xref ref-type="bibr" rid="CIT177">Jak+22</xref>] or PRDPT [<xref ref-type="bibr" rid="CIT092">FR22</xref>] for inverse rendering applications. All in all, NEnv can represent global illumination with higher generality, quality and computational efficiency than previous work, with fully differentiable components.</p>
<p>Our work could be extended in several ways. We require two networks, one for sampling and pdf evaluation, and another for compression. While this dual approach allows us to maximize their individual efficiency, a future research avenue is to build a single model which can solve the three tasks at the same time, to simplify our representation and possibly help with parameter reuse across models. However, fusing a sinusoidal MLP with a spline-based normalizing flow is not straightforward and requires a significant research effort that lies outside of the scope of this work. Recent work on architectural design [<xref ref-type="bibr" rid="CIT267">Meh+21</xref>] may provide cues on how this could be achieved. Further, we require to train new models for each input environment map. Building upon <target target-type="page" id="pges_207"/>recent work [<xref ref-type="bibr" rid="CIT105">GES22</xref>; <xref ref-type="bibr" rid="CIT315">Rai+22</xref>; <xref ref-type="bibr" rid="CIT101">GMX22</xref>; <xref ref-type="bibr" rid="CIT380">Szt+21</xref>; <xref ref-type="bibr" rid="CIT159">Hu+20</xref>], we could build a prior over environment maps, which should help reduce training times and solve inverse illumination problems, as in [<xref ref-type="bibr" rid="CIT241">Li+20</xref>; <xref ref-type="bibr" rid="CIT016">Azi+19</xref>; <xref ref-type="bibr" rid="CIT136">HHM22</xref>; <xref ref-type="bibr" rid="CIT439">YS19</xref>; <xref ref-type="bibr" rid="CIT420">Wan+21</xref>; <xref ref-type="bibr" rid="CIT440">YS21</xref>; <xref ref-type="bibr" rid="CIT372">Sri+20</xref>; <xref ref-type="bibr" rid="CIT106">Gar+19</xref>; <xref ref-type="bibr" rid="CIT421">Wan+23</xref>]. Achieving this without losing computational efficiency is a challenging future research problem. Besides, recent work on neural rendering [<xref ref-type="bibr" rid="CIT283">M&#x00FC;l+22</xref>; <xref ref-type="bibr" rid="CIT044">Che+22</xref>; <xref ref-type="bibr" rid="CIT414">Wan+22d</xref>; <xref ref-type="bibr" rid="CIT015">Att+22</xref>; <xref ref-type="bibr" rid="CIT252">Liu+22a</xref>] has shown that with carefully designed CUDA kernels or architectural representations, it is possible to achieve more efficient training and evaluation times. Integrating these ideas into our framework could help further increase its computational efficiency.</p>
<p>Finally, we believe that the ideas we present in this work can be used for other path tracing and physically based rendering problems, like aggregate scattering [<xref ref-type="bibr" rid="CIT034">Blu+16</xref>] or BxDF representations [<xref ref-type="bibr" rid="CIT380">Szt+21</xref>; <xref ref-type="bibr" rid="CIT050">CNN22</xref>; <xref ref-type="bibr" rid="CIT446">Zha+21</xref>; <xref ref-type="bibr" rid="CIT043">Che+20a</xref>]. We hope our work inspires future work on generative models for physically-based and neural rendering.</p>
</sec>
</body>
</book-part>
<book-part id="c8" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_208"/><target target-type="page" id="pges_209"/>Chapter 8.</label>
<title>Conclusions</title>
</title-group>
</book-part-meta>
<body>
<p>Throughout this thesis, we have addressed important challenges in the fields of digital materials and radiance encoding, including <italic>Visual Attribute Transfer</italic> (<xref ref-type="book-part" rid="c2">Chapter 2</xref>), <italic>SVBSDF Capture</italic> (<xref ref-type="sec" rid="c2-sC">Chapter 2.C</xref>), <italic>Tileable Texture Synthesis</italic> (<xref ref-type="book-part" rid="c3">Chapter 3</xref>), <italic>Neural BTF Propagation and Encoding</italic> (<xref ref-type="book-part" rid="c4">Chapter 4</xref>), <italic>SVBRDF Estimation</italic> (<xref ref-type="book-part" rid="c5">Chapter 5</xref>), <italic>Fabric Mechanical Parameter Capture</italic> (<xref ref-type="book-part" rid="c6">Chapter 6</xref>) and <italic>Neural Methods for Global Illumination</italic> (<xref ref-type="book-part" rid="c7">Chapter 7</xref>). In this final section, our goal is to provide a comprehensive summary of the contributions made in this thesis and their significance in the fields of computer vision, graphics, and machine learning. By providing a broader view, we aim to highlight the importance and impact of the work presented here, both in terms of advancing the state-of-the-art and in its potential for real-world applications.</p>
<sec id="c8-s1">
<title>Digital Material Generation</title>
<p>The first projects developed in the context of this thesis are concerned with generating high quality digital materials, which can then be used for other tasks, like training neural material representations or creating large datasets for material estimation models. These methods take as input a single material, which helps increase the predictability of their outputs.</p>
<p>In <xref ref-type="book-part" rid="c2">Chapter 2</xref>, we present a neural visual attribute transfer framework capable of transferring, for a given material, many types of visual property maps to images of new patches of the same (or similar) material. These novel images may be taken under any illumination conditions, color variations, affine distortions, or camera setup. This model uses lightweight image-to-image convolutional neural networks to provide the first solution capable of leveraging the optical behavior of the material for the visual attribute transfer problem, by being trained on with a photometric dataset and a comprehensive data augmentation policy. This method can be trained in less than one minute, and is capable of transferring any type of visual attribute, from surface normals to stylizations or semantic segmentation maps. However, it is limited in accuracy, can only learn from one dataset at once, and can only learn to transfer one visual attribute.</p>
<p>To address these shortcomings, we extend this method in <xref ref-type="sec" rid="c2-sC">Chapter 2.C</xref>, where we propose a neural network capable of generating high-resolution large-scale SVBSDF maps of any material given one or more photometric datasets which represent small portions of a material. By extensive changes to the training procedure and improvements to the model architecture, not only do we enable this visual attribute transfer with the capability to learn from multiple datasets and transfer multiple images simultaneously, but we also increase the accuracy <target target-type="page" id="pges_210"/>and sharpness of its estimations. However, these improvements increase the computational cost of training the neural network, and limit their generality. With these models, we can build large-scale materials, which can be used to generate BTF data to train other representations (<xref ref-type="book-part" rid="c4">Chapter 4</xref>) or build large datasets to train models that generalize to new materials (<xref ref-type="book-part" rid="c5">Chapter 5</xref>).</p>
<p>Textures used in rendering applications also need to be tileable. That is, when they are spatially concatenated, the seams between tiles are not visible. In <xref ref-type="book-part" rid="c3">Chapter 3</xref>, we introduce a single-image Generative Adversarial Network (GAN) which can generate multiple tileable textures of the same material. Our method allows for tileable outputs using a novel latent space manipulation algorithm, and handles multiple texture maps simultaneously, through neural architecture improvements. Further, we leverage the GAN&#x2019;s discriminator as a test-time quality estimation function, which helps distinguish between high quality and low quality outputs. This model achieves higher quality outputs than previous work, at lower computational times.</p>
<p>These models rely on a single material as input, which, while increasing predictability, limits their generalization capabilities. Nonetheless, our findings provide evidence that under certain assumptions and highly controlled training procedures, neural networks can be utilized for generating ground truth data. This data can be leveraged for other tasks, as demonstrated in other sections of this thesis.</p>
</sec>
<sec id="c8-s2">
<title>Neural Radiance Encoding</title>
<p>One of the research problems that we explore in this thesis is the challenge of representing radiance using neural networks, which offers several potential <target target-type="page" id="pges_211"/>benefits compared to traditional approaches. Neural networks can provide efficient, continuous, and fully differentiable functions that are useful for forward and inverse rendering. Additionally, they enable novel editing capabilities and efficient sampling for Monte Carlo rendering. In this work, we focus on two key components of virtual scenes: materials and illumination. Our goal is to develop neural network-based approaches for encoding these elements that are both accurate and computationally efficient, paving the way for new applications in computer graphics and vision.</p>
<p>For materials, in <xref ref-type="book-part" rid="c4">Chapter 4</xref>, we introduce a novel neural representation for encoding reflectance. Building upon previous work on the topic of neural materials [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>], our method has three components: A <italic>neural texture</italic>, which encodes the reflectance of the material in a 2D latent space; a fully-convolutional <italic>neural renderer</italic>, which takes as input this latent reflectance and light and view angles and outputs linear RGB values; and an <italic>autoencoder</italic>, which can estimate the neural texture of any input 2D image of the material. Similarly to previous work on neural fields [<xref ref-type="bibr" rid="CIT276">Mil+20</xref>], we train each model from scratch for every material, taking as input its <italic>Bidirectional Texture Function</italic> (BTF), which can be either captured or synthetically generated. Our approach introduces the first neural BTF representation with conditional input, which makes it possible to extrapolate BTF measurements, create tileable BTFs, and synthesize new materials through reflectance propagation. We address one of the main limitations of prior work on neural materials, which was their lack of editing capabilities, hindering their applicability to real-world rendering scenarios that require tileable and large-scale materials. Despite these improvements, our representation has limitations in that semantic properties such as specularity or albedo are encoded on the latent space and thus not easily editable. However, we believe that our method has significant potential for future research on neural material representations.</p>
<p>In <xref ref-type="book-part" rid="c7">Chapter 7</xref>, we addressed the problem of illumination encoding, introducing a novel neural method for joint sampling, PDF encoding, and RGB compression of environment maps used for global illumination in rendering. Our method builds on previous work on <italic>implicit neural representations</italic> for compressing the HDRi environment map into a continuous, differentiable function. We designed a <italic>Normalizing Flow</italic> that can sample directions from environment maps and measure their probability density, both of which are essential for Multiple Importance Sampling in Monte Carlo rendering. Our carefully designed neural architecture and training procedures allowed us to obtain models that are up to two orders of magnitude faster than analytical methods while still being fully differentiable. Using a dataset of environment maps with diverse properties, we demonstrated high generality and accuracy, outperforming previous work on the topic. To encourage future research, we will provide an open-source implementation and a dataset of trained models. While this is the first method that learns to sample from environment maps, we believe there is potential <target target-type="page" id="pges_212"/>to further improve our approach by introducing more sophisticated neural network designs or learning priors over global illumination.</p>
<p>Despite their shortcomings, these two neural methods introduce contributions to the field of neural scene representations, and they could both be incorporated into neural rendering pipelines and more traditional path-tracing engines. We hope these ideas inspire further research interest in this field.</p>
</sec>
<sec id="c8-s3">
<title>Scalable Material Digitization</title>
<p>One of the key objectives of this thesis is to develop digitization solutions that are both scalable and affordable, without relying on expensive or inaccessible hardware. To accomplish this, we have made two different contributions. First, in <xref ref-type="book-part" rid="c5">Chapter 5</xref>, we introduce a novel method for estimating <italic>spatially varying bidirectional reflectance distribution functions</italic> (SVBRDFs) using flatbed scanners as the capture device. This method strongly relies on the data generation algorithms presented in previous sections of this thesis. Our approach employs a custom-built, lightweight neural network with attention mechanisms, resulting in estimations with considerably higher resolution than previous methods that relied on smartphone flash-lit images as input. We empirically observe that material reflectance is strongly dependent on its microgeometry, which flatbed scanners can capture appropriately, all the while providing images that can be directly used as the albedo of the SVBRDF. One of the key contributions of this work is the introduction of the first uncertainty quantification algorithm for material estimation, for which we demonstrate applications on dataset creation through active learning. Further, an extension of this work is part of <uri xlink:href="http://www.Textura.ai">Textura.ai</uri>, which is utilized by real-world users, providing evidence of its robustness and reliability. In <xref ref-type="fig" rid="fig-8.1">Figure 8.1</xref>, we show some digital garments which use materials digitized with this technology. We are confident that the novelties introduced by this method will shape future research on the long-standing problem of single image material digitization. Although we are confident in the method&#x2019;s capabilities, it can be extended to estimate additional reflectance properties like anisotropy or transmittance, to generalize to other materials or devices, and to improve its accuracy even further.</p>
<fig id="fig-8.1">
<label>Figure 8.1:</label>
<caption><title>Renders generated with SEDDI&#x2019;s <uri xlink:href="http://www.Textura.ai">Textura.ai</uri> materials, estimated with methods derived from some of the contents in this thesis.</title></caption>
<alt-text>Renders generated with SEDDI&#x2019;s Textura.ai materials, estimated with methods derived from some of the contents in this thesis.</alt-text>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="TD-13-43-URJC_fig-8.1.jpg"/>
</fig>
<p>Besides, in <xref ref-type="book-part" rid="c6">Chapter 6</xref>, we proposed a method for mechanical parameter estimation of textile materials using depth images and the material density as inputs. Our method achieves greater scalability than previous work, which required video sequences, manual intervention, or purposefully built and expensive devices. To do so, we propose a solution that is agnostic to the optical appearance of the fabric, and can be utilized by non-expert operators. Our custom-built multi-input neural network predicts the full set of mechanical parameters of any textile using synthetic data, leveraging a strong data <target target-type="page" id="pges_213"/>augmentation policy, pre-training, attention mechanisms, and multiple forms of pooling of internal activations. A significant contribution of this work is our proposed <italic>drape similarity metric</italic>, which measures distances between the mechanical behavior of fabrics while using only images as input. We validatethis metric with a user study conducted on-site, showing that it strongly correlates with human perception. Although our method has significant contributions to estimating the mechanical behavior of textiles, it is still limited in at least two ways: it takes depth images as input, which are relatively easy to obtain but are less accessible than RGB pictures, and the proposed drape similarity metric is not easily differentiable, hindering its potential as a loss function in optimization problems.</p>
<p>These two estimation algorithms, along with the uncertainty quantification algorithm and the proposed drape similarity metric, have the potential to transform the low cost material digitization field, introducing more scalable, robust, and perceptually-validated solutions.</p>
</sec>
<sec id="c8-s4">
<title>New Frontiers</title>
<p>Although this thesis has made significant contributions to the fields of radiance encoding, and digital material capture and edition, we acknowledge that our proposed methods have certain limitations in terms of scope and generality. Additionally, concurrent developments in deep learning, neural rendering, and computer vision have opened up new and exciting research avenues that will shape the future of the field. With this in mind, we would like to highlight several research directions that we believe are particularly intriguing for extending the work presented in this thesis.</p>
</sec>
<sec id="c8-s5">
<title>Generative Models for Materials</title>
<p>In this thesis, we have presented models that are limited to solving individual tasks such as material digitization, encoding, or tileable texture synthesis. However, recent research [<xref ref-type="bibr" rid="CIT343">Rom+22</xref>; <xref ref-type="bibr" rid="CIT157">HJA20</xref>; <xref ref-type="bibr" rid="CIT349">Sah+22</xref>] has demonstrated that a single generative model can be used to solve multiple tasks simultaneously. Developing a single generative model for materials that can be used for capture, super-resolution, edition, real-time tileable synthesis, or material interpolation would be an exciting research direction for the future of digital materials. One of the main challenges in this area is the lack of publicly available datasets of real materials, which are essential for reproducible research on deep learning algorithms. Nevertheless, techniques like pre-training, fine-tuning, and data augmentation policies may be employed to construct these models without requiring massive datasets of materials.</p>
</sec>
<sec id="c8-s6">
<title><target target-type="page" id="pges_214"/>Multimodal Learning</title>
<p>The findings in this thesis suggest an interesting research avenue: multimodal learning for materials. In addition to capturing a material&#x2019;s reflectance and mechanical properties, it is possible to digitize other attributes such as its fabrication process, semantic attributes, fine-grained composition, origin, cost, or even its carbon footprint. Recent work on multimodal and contrastive learning [<xref ref-type="bibr" rid="CIT303">Par+20</xref>] suggests that a shared embedding for visual and semantic attributes can be found, which can then be utilized for various tasks such as zero-shot learning or conditioning generative models [<xref ref-type="bibr" rid="CIT343">Rom+22</xref>]. Developing such an embedding for materials could be extremely valuable for a wide range of problems, including material edition, interpolation, classification, generation, similarity measurement, or anomaly detection. This research direction is promising, but it also poses significant challenges, such as the need for large and diverse datasets that include multiple modalities of material information, and the development of effective multimodal learning architectures that can handle such information. Nonetheless, we believe that the benefits of multimodal learning for materials make it a hugely compelling direction for future research.</p>
</sec>
<sec id="c8-s7">
<title>Neural Spatially-Varying BSDFs</title>
<p>In <xref ref-type="book-part" rid="c7">Chapter 7</xref>, we demonstrated the effectiveness of invertible generative neural networks for efficient Multiple Importance Sampling in global illumination. Similar networks have also been utilized for BRDFs [<xref ref-type="bibr" rid="CIT380">Szt+21</xref>], suggesting that learned sampling could be applied to spatially-varying materials. Further, in <xref ref-type="book-part" rid="c4">Chapter 4</xref>, we propose a neural representation of spatially-varying materials, which showed promising compression results. Combining the ideas of these two projects, it should be possible to train invertible neural networks to learn to sample from real-world materials, which often exhibit complex optical behavior such as spatially-varying reflectance or transmittance, we can potentially improve the efficiency and quality of path-traced renders while enabling new inverse rendering applications. This presents a promising direction for future research in material capture and rendering.</p>
</sec>
<sec id="c8-s8">
<title>Uncertainty Quantification and Model Understanding</title>
<p>We have presented the first method for uncertainty quantification in material capture applications, demonstrating its utility for active learning and dataset creation. Additionally, we have shown that visualization of material capture models can aid in understanding their predictions. However, we believe that these areas are still relatively unexplored, and further analysis could provide valuable insights for designing the next generation of material digitization <target target-type="page" id="pges_215"/>pipelines. This includes developing more suitable capture devices, identifying valuable data points for labeling, and building more accurate models. Furthermore, we believe that investigating methods for communicating to end users when a capture may be inaccurate and suggesting alternative digitization pathways is an important research topic that could enhance the impact, robustness, and trustworthiness of these models in the industry.</p>
</sec>
<sec id="c8-s9">
<title>Enhancing Digitization Quality and Realism</title>
<p>In this thesis, we have presented methods for highly realistic material encodings (<xref ref-type="book-part" rid="c4">Chapter 4</xref>) and digitization pipelines (<xref ref-type="sec" rid="c2-sC">Chapter 2.C</xref>). They can be used to represent a wide variety of materials at very high qualities, however, they require expensive capture systems. On the other hand, in <xref ref-type="book-part" rid="c5">Chapter 5</xref>, we introduce a single-image digitization system capable of capturing a small set of SVBRDF parameters. While the model is adequately accurate, this limited material model cannot represent complex optical phenomena, like anisotropy or transmittance. Extending single-image material digitization systems to be able to accurately estimate more reflectance properties is a very challenging research problem, because this estimation is inherently ill-posed, and there is very limited available training data. These low cost capture systems will be able to increase the realism of their digitizations in the future, as more and better data becomes available, and better models can be built as the machine learning field advances.</p>
</sec>
<sec id="c8-s10">
<title>Enhancing Generality on Devices and Materials</title>
<p>A limitation of our casual capture algorithms is that they require the use of specific, albeit low-cost, devices such as flatbed scanners or depth cameras. An immediate extension of our methods is to make them compatible with other types of devices, particularly smartphone cameras, which would increase their applicability. However, less controlled capture settings typically hinder the accuracy of digitizations, as additional assumptions must be made to obtain a plausible result. With the ongoing improvement in the quality of smartphone cameras, along with larger datasets and better deep learning models, high-quality digitization of material reflectance or fabric mechanics using smartphones could become a reality in the near future.</p>
<p>In addition, like any learning-based system, our models are constrained by the materials present in their training set, limiting their ability to generalize to other materials. Although we have achieved accurate digitization of fabrics and leathers, there is no guarantee that our models will perform well on out-of-distribution materials such as stone, plastics, metals, or glass. While some of these materials may be successfully digitized using single-image <target target-type="page" id="pges_216"/>models with sufficient data, others may present more significant challenges. It is our hope that future research in this area will help improve the generality and accuracy of these models.</p>
</sec>
<sec id="c8-s11">
<title>Digitization on Edge Devices</title>
<p>One limitation of material capture algorithms based on deep learning is that they require expensive hardware for inference. This means that in commercial products, users can capture images on a commodity device, but the digitization process is typically carried out on a cloud-based application. As a result, bottlenecks may arise that hinder the efficiency of the digitization process, with users having to wait several minutes before obtaining their digital material. To achieve real-time learning-based digitization, it may be necessary to perform model inference on edge devices such as smartphones, tablets, or laptops. However, this approach may place limitations on model sizes, potentially leading to a reduction in digitization accuracy. Recent work on edge deep learning [<xref ref-type="bibr" rid="CIT417">Wan+20c</xref>] may provide interesting cues for future research on this topic.</p>
</sec>
<sec id="c8-s12">
<title>Neural Scene Representations</title>
<p>Concurrently to the work in this thesis, <italic>Neural Radiance Fields</italic> [<xref ref-type="bibr" rid="CIT276">Mil+20</xref>] emerged as an alternative powerful representation for virtual scenes. A Cambrian explosion of research in radiance fields for scene representations has emerged to increase their capabilities, in terms of data efficiency, speed, realism, or generality [<xref ref-type="bibr" rid="CIT414">Wan+22d</xref>; M&#x00FC;l+22; Wan+23; Che+22; Tew+22; Ver+22; Fri+22; Att+22; Xie+22]. However, these remain somewhat limited in controllable and capabilities for edition, for which more traditional rendering and scene creation pipelines may be more adequate than learning-based alternatives. Our neural encodings for illumination or materials, which may be integrated into either neural representations or path-traced rendering engines, may help bridge these two pipelines, which should benefit both.</p>
</sec>
<sec id="c8-s13">
<title>Dataset Generation for Vision Downstream Tasks</title>
<p>Synthetic data has been extensively shown to help deep learning models generalize to real images [<xref ref-type="bibr" rid="CIT017">Azi+23</xref>; <xref ref-type="bibr" rid="CIT462">Zhu+21</xref>; <xref ref-type="bibr" rid="CIT335">Ric+16</xref>]. Our proposed method for casual capture of mechanical parameters in fabrics (<xref ref-type="book-part" rid="c6">Chapter 6</xref>, [<xref ref-type="bibr" rid="CIT340">Rod+23b</xref>]) provides additional evidence in this direction. The procedural creation of virtual environments for pre-training of computer vision models will remain a very relevant topic in future years in the field, as it will reduce the need for expensive labeling of real-world data. These virtual environments strongly <target target-type="page" id="pges_217"/>benefit from highly accurate assets for enhanced realism, as well as efficient yet expressive image formation algorithms. We hope that the contributions in this thesis will help in generating synthetic data for training computer vision models which solve general tasks, including classification or segmentation [<xref ref-type="bibr" rid="CIT208">Kir+23</xref>].</p>
</sec>
<sec id="c8-s14">
<title>Applications to Other Fields</title>
<p>Finally, a fascinating use of computer vision and material digitization is <italic>digital cultural heritage</italic> [<xref ref-type="bibr" rid="CIT130">HMV09</xref>]. Digital capture of artifacts from ancient civilizations could help important research fields, like anthropology or archaeology. However, past solutions required inaccessible devices [<xref ref-type="bibr" rid="CIT132">Ham+21</xref>; <xref ref-type="bibr" rid="CIT082">Dye+18</xref>], which limited their impact. Inexpensive gathering of photorealistic copies of ancient artifacts, which can be shared digitally, can provide valuable solutions for these research fields. We hope that low cost and accurate digitization systems may not only help the visual media and design industries, but also other fields for which digital asset inventories are helpful.</p>
</sec>
<sec id="c8-s15">
<title>Closing Remarks</title>
<p>The collaborative efforts between disciplines and research institutions, both public and industrial, have been instrumental in enabling the contributions of this thesis, and the advancements made in the computer vision, graphics, and machine learning fields as a whole in the past few years. The interdisciplinary environment in which I worked during my PhD greatly enhanced the quality of my research and communication skills, while also providing me with the opportunity to learn fascinating new topics and collaborate with exceptional people. Further, while the massive amounts of publications in these fields put a significant amount of pressure on the researchers, it makes the field move incredibly quickly and creates a very exciting environment to work in. I hope that the focus on open science and the integration of ideas and feedback from other research areas continues, along with a greater awareness of the ethical and ecological impact of the models we create and deploy. Working as a researcher is in many ways a privilege, and it is our responsibility to advance science in a way that positively impacts society and the environment, and creates a better, more prosperous, sustainable, and equitable future for everyone. Looking ahead to the future, my hope is that the contributions made in this thesis will have a tangible impact on the real world and serve to further advance research in positive ways. On a personal level, my goal is to continue learning and challenging myself, all while striving to make a positive difference in the world.</p>
</sec>
</body>
</book-part>
</book-body>
<book-back>
<book-part id="c9" book-part-type="chapter">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<title><target target-type="page" id="pges_218"/><target target-type="page" id="pges_219"/>Bibliography</title>
</title-group>
</book-part-meta>
<back>
<ref-list id="bib01">
<ref id="CIT001"><label>[Aba+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Mart&#x0131;n</given-names> <surname>Abadi</surname></string-name>, <string-name><given-names>Paul</given-names> <surname>Barham</surname></string-name>, <string-name><given-names>Jianmin</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Zhifeng</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Andy</given-names> <surname>Davis</surname></string-name>, <string-name><given-names>Jeffrey</given-names> <surname>Dean</surname></string-name>, <string-name><given-names>Matthieu</given-names> <surname>Devin</surname></string-name>, <string-name><given-names>Sanjay</given-names> <surname>Ghemawat</surname></string-name>, <string-name><given-names>Geoffrey</given-names> <surname>Irving</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Isard</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;Tensorflow: A System for Large-Scale Machine Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>12th Symposium on Operating Systems Design and Implementation</italic></source>. <year>2016</year>, <comment>pp.</comment> <fpage>265</fpage>&#x2013;<lpage>283</lpage> (<comment>cit. on p. 96</comment>).</mixed-citation></ref>
<ref id="CIT002"><label>[Aga+03]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sameer</given-names> <surname>Agarwal</surname></string-name>, <string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name>, <string-name><given-names>Serge</given-names> <surname>Belongie</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Henrik</given-names> <surname>Wann Jensen</surname></string-name></person-group>. <article-title>&#x201C;Structured Importance Sampling of Environment Maps&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>22</volume>.<issue>3</issue> (<year>2003</year>), <comment>pp.</comment> <fpage>605</fpage>&#x2013;<lpage>612</lpage> (<comment>cit. on p. 185</comment>).</mixed-citation></ref>
<ref id="CIT003"><label>[AAL16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Miika</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>Timo</given-names> <surname>Aila</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jaakko</given-names> <surname>Lehtinen</surname></string-name></person-group>. <article-title>&#x201C;Reflectance Modeling by Neural Texture Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>35</volume>.<issue>4</issue> (<year>2016</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>13</lpage> (<comment>cit. on p. 124</comment>).</mixed-citation></ref>
<ref id="CIT004"><label>[AWL13]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Miika</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>Tim</given-names> <surname>Weyrich</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jaakko</given-names> <surname>Lehtinen</surname></string-name></person-group>. <article-title>&#x201C;Practical SVBRDF Capture in the Frequency Domain&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>32</volume>.<issue>4</issue> (<year>2013</year>), <comment>pp.</comment> <fpage>110</fpage>&#x2013;<lpage>1</lpage> (<comment>cit. on pp. 17, 107, 128, 258</comment>).</mixed-citation></ref>
<ref id="CIT005"><label>[AWL15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Miika</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>Tim</given-names> <surname>Weyrich</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jaakko</given-names> <surname>Lehtinen</surname></string-name></person-group>. <article-title>&#x201C;Two-shot SVBRDF Capture for Stationary Materials&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>34</volume>.<issue>4</issue> (<year>2015</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>13</lpage> (<comment>cit. on pp. 17, 34, 104, 125, 258</comment>).</mixed-citation></ref>
<ref id="CIT006"><label>[Akl+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Adib</given-names> <surname>Akl</surname></string-name>, <string-name><given-names>Charles</given-names> <surname>Yaacoub</surname></string-name>, <string-name><given-names>Marc</given-names> <surname>Donias</surname></string-name>, <string-name><given-names>Jean-Pierre</given-names> <surname>Da Costa</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Christian</given-names> <surname>Germain</surname></string-name></person-group>. <article-title>&#x201C;A Survey of Exemplar-Based Texture Synthesis Methods&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Vision and Image Understanding</italic></source> <volume>172</volume> (<year>2018</year>), <comment>pp.</comment> <fpage>12</fpage>&#x2013;<lpage>24</lpage> (<comment>cit. on p. 73</comment>).</mixed-citation></ref>
<ref id="CIT007"><label>[Alc+19]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Raul</given-names> <surname>Alcain</surname></string-name>, <string-name><given-names>Carlos</given-names> <surname>Heras</surname></string-name>, <string-name><given-names>I&#x00F1;igo</given-names> <surname>Salinas</surname></string-name>, <string-name><given-names>Jorge</given-names> <surname>Lopez-Moreno</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Aliaga</surname></string-name></person-group>. <chapter-title>&#x201C;Microscale Optical Capture System for Digital Fabric Recreation&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the 7th International Conference on Photonics, Optics and Laser Technology - Volume 1: PHOTOPTICS,</italic></source> <publisher-name>INSTICC. SciTePress</publisher-name>, <year>2019</year>, <comment>pp.</comment> <fpage>114</fpage>&#x2013;<lpage>119</lpage> (<comment>cit. on pp. 35, 39</comment>).</mixed-citation></ref>
<ref id="CIT008"><label>[Alm+18]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Amjad</given-names> <surname>Almahairi</surname></string-name>, <string-name><given-names>Sai</given-names> <surname>Rajeshwar</surname></string-name>, <string-name><given-names>Alessandro</given-names> <surname>Sordoni</surname></string-name>, <string-name><given-names>Philip</given-names> <surname>Bachman</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Aaron</given-names> <surname>Courville</surname></string-name></person-group>. <chapter-title>&#x201C;Augmented Cyclegan: Learning Many-to-Many Mappings from Unpaired Data&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2018</year>, <comment>pp.</comment> <fpage>195</fpage>&#x2013;<lpage>204</lpage> (<comment>cit. on p. 85</comment>).</mixed-citation></ref>
<ref id="CIT009"><label>[Ami+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexander</given-names> <surname>Amini</surname></string-name>, <string-name><given-names>Wilko</given-names> <surname>Schwarting</surname></string-name>, <string-name><given-names>Ava</given-names> <surname>Soleimany</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Daniela</given-names> <surname>Rus</surname></string-name></person-group>. <article-title>&#x201C;Deep Evidential Regression&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in <target target-type="page" id="pges_220"/>Neural Information Processing Systems</italic></source> <lpage>33</lpage> (<year>2020</year>), <fpage>pp.</fpage> <fpage>14927</fpage>&#x2013;<lpage>14937</lpage> (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT010"><label>[AP08]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xiaobo</given-names> <surname>An</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Fabio</given-names> <surname>Pellacini</surname></string-name></person-group>. <article-title>&#x201C;AppProp: All-pairs Appearance-Space Edit Propagation&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>27</volume>.<issue>3</issue> (<year>2008</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>9</lpage> (<comment>cit. on p. 33</comment>).</mixed-citation></ref>
<ref id="CIT011"><label>[And+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Pontus</given-names> <surname>Andersson</surname></string-name>, <string-name><given-names>Jim</given-names> <surname>Nilsson</surname></string-name>, <string-name><given-names>Tomas</given-names> <surname>Akenine-M&#x00F6;ller</surname></string-name>, <string-name><given-names>Magnus</given-names> <surname>Oskarsson</surname></string-name>, <string-name><given-names>Kalle</given-names> <surname>&#x00C5;str&#x00F6;m</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mark D</given-names> <surname>Fairchild</surname></string-name></person-group>. <article-title>&#x201C;FLIP: A Difference Evaluator for Alternating Images.&#x201D;</article-title> <comment>In</comment>: <source><italic>Proc. ACM Comput. Graph. Interact. Tech</italic></source>. <volume>3</volume>.<issue>2</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>15</fpage>&#x2013;<lpage>1</lpage> (<comment>cit. on pp. 198, 199, 202</comment>).</mixed-citation></ref>
<ref id="CIT012"><label>[ASE17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Antreas</given-names> <surname>Antoniou</surname></string-name>, <string-name><given-names>Amos</given-names> <surname>Storkey</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Harrison</given-names> <surname>Edwards</surname></string-name></person-group>. <article-title>&#x201C;Data Augmentation Generative Adversarial Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1711.04340</italic></source> (<year>2017</year>) (<comment>cit. on p. 45</comment>).</mixed-citation></ref>
<ref id="CIT013"><label>[ACB19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Devansh</given-names> <surname>Arpit</surname></string-name>, <string-name><given-names>V&#x0131;ctor</given-names> <surname>Campos</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yoshua</given-names> <surname>Bengio</surname></string-name></person-group>. <article-title>&#x201C;How to Initialize Your Network? Robust Initialization for Weightnorm &#x0026; Resnets&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>32</volume> (<year>2019</year>) (<comment>cit. on pp. 109, 119</comment>).</mixed-citation></ref>
<ref id="CIT014"><label>[ARV19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yuki M</given-names> <surname>Asano</surname></string-name>, <string-name><given-names>Christian</given-names> <surname>Rupprecht</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Andrea</given-names> <surname>Vedaldi</surname></string-name></person-group>. <article-title>&#x201C;A Critical Analysis of Self-Supervision, or What We Can Learn from a Single Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2019</year> (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT015"><label>[Att+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Benjamin</given-names> <surname>Attal</surname></string-name>, <string-name><given-names>Jia-Bin</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Zollh&#x00F6;fer</surname></string-name>, <string-name><given-names>Johannes</given-names> <surname>Kopf</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Changil</given-names> <surname>Kim</surname></string-name></person-group>. <article-title>&#x201C;Learning Neural Light Fields with Ray-space Embedding&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <fpage>pp.</fpage> <fpage>19819</fpage>&#x2013;<lpage>19829</lpage> (<comment>cit. on pp. 187, 203, 212</comment>).</mixed-citation></ref>
<ref id="CIT016"><label>[Azi+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dejan</given-names> <surname>Azinovic</surname></string-name>, <string-name><given-names>Tzu-Mao</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Anton</given-names> <surname>Kaplanyan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Matthias</given-names> <surname>Nie&#x00DF;ner</surname></string-name></person-group>. <article-title>&#x201C;Inverse Path Tracing for Joint Material and Lighting Estimation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>2447</fpage>&#x2013;<lpage>2456</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT017"><label>[Azi+23]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Shekoofeh</given-names> <surname>Azizi</surname></string-name>, <string-name><given-names>Simon</given-names> <surname>Kornblith</surname></string-name>, <string-name><given-names>Chitwan</given-names> <surname>Saharia</surname></string-name>, <string-name><given-names>Mohammad</given-names> <surname>Norouzi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>David J</given-names> <surname>Fleet</surname></string-name></person-group>. <article-title>&#x201C;Synthetic Data from Diffusion Models Improves ImageNet Classification&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2304.08466</italic></source> (<year>2023</year>) (<comment>cit. on p. 212</comment>).</mixed-citation></ref>
<ref id="CIT018"><label>[BKH16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jimmy Lei</given-names> <surname>Ba</surname></string-name>, <string-name><given-names>Jamie Ryan</given-names> <surname>Kiros</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Geoffrey E</given-names> <surname>Hinton</surname></string-name></person-group>. <article-title>&#x201C;Layer Normalization&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1607.06450</italic></source> (<year>2016</year>) <comment>(cit. on pp. 109, 111, 119, 176</comment>).</mixed-citation></ref>
<ref id="CIT019"><label>[Baa+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Hendrik</given-names> <surname>Baatz</surname></string-name>, <string-name><given-names>Jonathan</given-names> <surname>Granskog</surname></string-name>, <string-name><given-names>Marios</given-names> <surname>Papas</surname></string-name>, <string-name><given-names>Fabrice</given-names> <surname>Rousselle</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Nov&#x00E1;k</surname></string-name></person-group>. <article-title>&#x201C;NeRF-Tex: Neural Reflectance <target target-type="page" id="pges_221"/>Field Textures&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 41</volume>. <issue>6</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2022</year>, <comment>pp.</comment> <fpage>287</fpage>&#x2013;<lpage>301</lpage> (<comment>cit. on pp. 104, 187</comment>).</mixed-citation></ref>
<ref id="CIT020"><label>[BSK23]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Steve</given-names> <surname>Bako</surname></string-name>, <string-name><given-names>Pradeep</given-names> <surname>Sen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Anton</given-names> <surname>Kaplanyan</surname></string-name></person-group>. <article-title>&#x201C;Deep Appearance Prefiltering&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>42</volume>.<issue>2</issue> (<year>2023</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>23</lpage> (<comment>cit. on p. 105</comment>).</mixed-citation></ref>
<ref id="CIT021"><label>[BYK21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wentao</given-names> <surname>Bao</surname></string-name>, <string-name><given-names>Qi</given-names> <surname>Yu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yu</given-names> <surname>Kong</surname></string-name></person-group>. <article-title>&#x201C;Evidential Deep Learning for Open Set Action Recognition&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>13349</fpage>&#x2013;<lpage>13358</lpage> (<comment>cit. on p. 130</comment>).</mixed-citation></ref>
<ref id="CIT022"><label>[Bar+09]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Connelly</given-names> <surname>Barnes</surname></string-name>, <string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name>, <string-name><given-names>Adam</given-names> <surname>Finkelstein</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Dan B</given-names> <surname>Goldman</surname></string-name></person-group>. <article-title>&#x201C;PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>28</volume>.<issue>3</issue> (<year>2009</year>), <comment>p.</comment> <fpage>24</fpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT023"><label>[Bar+15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Connelly</given-names> <surname>Barnes</surname></string-name>, <string-name><given-names>Fang-Lue</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Liming</given-names> <surname>Lou</surname></string-name>, <string-name><given-names>Xian</given-names> <surname>Wu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Shi-Min</given-names> <surname>Hu</surname></string-name></person-group>. <article-title>&#x201C;Patchtable: Efficient Patch Queries for Large Datasets and Applications&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>34</volume>.<issue>4</issue> (<year>2015</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>10</lpage> (<comment>cit. on pp. 32, 74</comment>).</mixed-citation></ref>
<ref id="CIT024"><label>[Bar+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jonathan T.</given-names> <surname>Barron</surname></string-name>, <string-name><given-names>Ben</given-names> <surname>Mildenhall</surname></string-name>, <string-name><given-names>Matthew</given-names> <surname>Tancik</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Hedman</surname></string-name>, <string-name><given-names>Ricardo</given-names> <surname>Martin-Brualla</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Pratul P.</given-names> <surname>Srinivasan</surname></string-name></person-group>. <article-title>&#x201C;Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT025"><label>[Ben+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Saguy</given-names> <surname>Benaim</surname></string-name>, <string-name><given-names>Ron</given-names> <surname>Mokady</surname></string-name>, <string-name><given-names>Amit</given-names> <surname>Bermano</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Lior</given-names> <surname>Wolf</surname></string-name></person-group>. <article-title>&#x201C;Structural Analogy from a Single Image Pair&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 40</volume>. <issue>1</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>249</fpage>&#x2013;<lpage>265</lpage> (<comment>cit. on pp. 31, 33, 46, 47, 56, 57, 60, 75</comment>).</mixed-citation></ref>
<ref id="CIT026"><label>[B&#x00E9;n+13]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Pierre</given-names> <surname>B&#x00E9;nard</surname></string-name>, <string-name><given-names>Forrester</given-names> <surname>Cole</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Kass</surname></string-name>, <string-name><given-names>Igor</given-names> <surname>Mordatch</surname></string-name>, <string-name><given-names>James</given-names> <surname>Hegarty</surname></string-name>, <string-name><given-names>Martin</given-names> <surname>Sebastian Senn</surname></string-name>, <string-name><given-names>Kurt</given-names> <surname>Fleischer</surname></string-name>, <string-name><given-names>Davide</given-names> <surname>Pesare</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Katherine</given-names> <surname>Breeden</surname></string-name></person-group>. <article-title>&#x201C;Stylizing Animation by Example&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>32</volume>.<issue>4</issue> (<year>2013</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>12</lpage> (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT027"><label>[Ben+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Gregory</given-names> <surname>Benton</surname></string-name>, <string-name><given-names>Marc</given-names> <surname>Finzi</surname></string-name>, <string-name><given-names>Pavel</given-names> <surname>Izmailov</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Andrew</given-names> <surname>Gordon Wilson</surname></string-name></person-group>. <article-title>&#x201C;Learning Invariances in Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2010.11882</italic></source> (<year>2020</year>) (<comment>cit. on p. 43</comment>).</mixed-citation></ref>
<ref id="CIT028"><label>[BJV17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Urs</given-names> <surname>Bergmann</surname></string-name>, <string-name><given-names>Nikolay</given-names> <surname>Jetchev</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Roland</given-names> <surname>Vollgraf</surname></string-name></person-group>. <article-title>&#x201C;Learning Texture Manifolds with the Periodic Spatial GAN&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>469</fpage>&#x2013;<lpage>477</lpage> (<comment>cit. on pp. 72, 74&#x2013;76, 89&#x2013;91, 95, 99</comment>).</mixed-citation></ref>
<ref id="CIT029"><label>[Ber+23]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Hugo</given-names> <surname>Bertiche</surname></string-name>, <string-name><given-names>Niloy J</given-names> <surname>Mitra</surname></string-name>, <string-name><given-names>Kuldeep</given-names> <surname>Kulkarni</surname></string-name>, <string-name><given-names>Chun-Hao</given-names> <surname>Paul Huang</surname></string-name>, <string-name><given-names>Tuanfeng Y</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Meysam</given-names> <surname>Madadi</surname></string-name>, <string-name><given-names>Sergio</given-names> <surname>Escalera</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Duygu</given-names> <surname>Ceylan</surname></string-name></person-group>. <target target-type="page" id="pges_222"/><article-title>&#x201C;Blowing in the Wind: CycleNet for Human Cinemagraphs from Still Images&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2303.08639</italic></source> (<year>2023</year>) (<comment>cit. on pp. 17, 259</comment>).</mixed-citation></ref>
<ref id="CIT030"><label>[Bha+03]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Kiran S.</given-names> <surname>Bhat</surname></string-name>, <string-name><given-names>Christopher D.</given-names> <surname>Twigg</surname></string-name>, <string-name><given-names>Jessica K.</given-names> <surname>Hodgins</surname></string-name>, <string-name><given-names>Pradeep K.</given-names> <surname>Khosla</surname></string-name>, <string-name><given-names>Zoran</given-names> <surname>Popovic</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Steven M.</given-names> <surname>Seitz</surname></string-name></person-group>. <chapter-title>&#x201C;Estimating Cloth Simulation Parameters from Video&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Symposium on Computer Animation</italic></source>. <publisher-name>The Eurographics Association</publisher-name>, <year>2003</year> (<comment>cit. on pp. 156, 158</comment>).</mixed-citation></ref>
<ref id="CIT031"><label>[Bi+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wenyan</given-names> <surname>Bi</surname></string-name>, <string-name><given-names>Peiran</given-names> <surname>Jin</surname></string-name>, <string-name><given-names>Hendrikje</given-names> <surname>Nienborg</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Bei</given-names> <surname>Xiao</surname></string-name></person-group>. <article-title>&#x201C;Estimating Mechanical Properties of Cloth from Videos Using Dense Motion Trajectories: Human Psychophysics and Machine Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>Journal of vision</italic></source> <volume>18</volume>.<issue>5</issue> (<year>2018</year>), <comment>pp.</comment> <fpage>12</fpage>&#x2013;<lpage>12</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT032"><label>[Bie20]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Lukas</given-names> <surname>Biewald</surname></string-name></person-group>. <source><italic>Experiment Tracking with Weights and Biases</italic></source>. <publisher-name>Software available from wandb.com</publisher-name>. <year>2020</year> (<comment>cit. on pp. 108, 176, 194</comment>).</mixed-citation></ref>
<ref id="CIT033"><label>[Bit+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Benedikt</given-names> <surname>Bitterli</surname></string-name>, <string-name><given-names>Chris</given-names> <surname>Wyman</surname></string-name>, <string-name><given-names>Matt</given-names> <surname>Pharr</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Shirley</surname></string-name>, <string-name><given-names>Aaron</given-names> <surname>Lefohn</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Wojciech</given-names> <surname>Jarosz</surname></string-name></person-group>. <article-title>&#x201C;Spatiotemporal Reservoir Resampling for Real-time Ray Tracing with Dynamic Direct Lighting&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (Proceedings of SIGGRAPH)</italic></source> <volume>39</volume>.<issue>4</issue> (<year>2020</year>) (<comment>cit. on p. 184</comment>).</mixed-citation></ref>
<ref id="CIT034"><label>[Blu+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Adrian</given-names> <surname>Blumer</surname></string-name>, <string-name><given-names>Jan</given-names> <surname>Nov&#x00E1;k</surname></string-name>, <string-name><given-names>Ralf</given-names> <surname>Habel</surname></string-name>, <string-name><given-names>Derek</given-names> <surname>Nowrouzezahrai</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Wojciech</given-names> <surname>Jarosz</surname></string-name></person-group>. <article-title>&#x201C;Reduced Aggregate Scattering Operators for Path Tracing&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum (Proceedings of Pacific Graphics)</italic></source> <volume>35</volume>.<issue>7</issue> (<year>2016</year>), <comment>pp.</comment> <fpage>461</fpage>&#x2013;<lpage>473</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT035"><label>[Bom+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Rishi</given-names> <surname>Bommasani</surname></string-name>, <string-name><given-names>Drew A</given-names> <surname>Hudson</surname></string-name>, <string-name><given-names>Ehsan</given-names> <surname>Adeli</surname></string-name>, <string-name><given-names>Russ</given-names> <surname>Altman</surname></string-name>, <string-name><given-names>Simran</given-names> <surname>Arora</surname></string-name>, <string-name><given-names>Sydney von</given-names> <surname>Arx</surname></string-name>, <string-name><given-names>Michael S</given-names> <surname>Bernstein</surname></string-name>, <string-name><given-names>Jeannette</given-names> <surname>Bohg</surname></string-name>, <string-name><given-names>Antoine</given-names> <surname>Bosselut</surname></string-name>, <string-name><given-names>Emma</given-names> <surname>Brunskill</surname></string-name>,</person-group> <etal>et al.</etal> <article-title>&#x201C;On the Opportunities and Risks of Foundation Models&#x201D;</article-title>. <comment>In</comment>: (<year>2021</year>) (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT036"><label>[Bos+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Mark</given-names> <surname>Boss</surname></string-name>, <string-name><given-names>Varun</given-names> <surname>Jampani</surname></string-name>, <string-name><given-names>Kihwan</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>Hendrik</given-names> <surname>Lensch</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Kautz</surname></string-name></person-group>. <article-title>&#x201C;Two-shot Spatially-Varying BRDF and Shape Estimation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>3982</fpage>&#x2013;<lpage>3991</lpage> (<comment>cit. on p. 125</comment>).</mixed-citation></ref>
<ref id="CIT037"><label>[Bou+13]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Katherine L</given-names> <surname>Bouman</surname></string-name>, <string-name><given-names>Bei</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Battaglia</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>William T</given-names> <surname>Freeman</surname></string-name></person-group>. <article-title>&#x201C;Estimating the Material Properties of Fabric from Video&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE International Conference on Computer Vision (ICCV)</italic></source>. <year>2013</year>, <comment>pp.</comment> <fpage>1984</fpage>&#x2013;<lpage>1991</lpage> (<comment>cit. on pp. 156, 159</comment>).</mixed-citation></ref>
<ref id="CIT038"><label>[BS12]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_223"/><person-group person-group-type="author"><string-name><given-names>Brent</given-names> <surname>Burley</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Walt</given-names> <surname>Disney Animation Studios</surname></string-name></person-group>. <article-title>&#x201C;Physically-based Shading at Disney&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM SIGGRAPH</italic></source>. <volume>Vol. 2012</volume>. <comment>vol. 2012</comment>. <year>2012</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>7</lpage> (<comment>cit. on pp. 76, 127</comment>).</mixed-citation></ref>
<ref id="CIT039"><label>[Byl+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zoya</given-names> <surname>Bylinskii</surname></string-name>, <string-name><given-names>Laura</given-names> <surname>Herman</surname></string-name>, <string-name><given-names>Aaron</given-names> <surname>Hertzmann</surname></string-name>, <string-name><given-names>Stefanie</given-names> <surname>Hutka</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yile</given-names> <surname>Zhang</surname></string-name></person-group>. <article-title>&#x201C;Towards Better User Studies in Computer Graphics and Vision&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2206.11461</italic></source> (<year>2022</year>) (<comment>cit. on pp. 178, 182</comment>).</mixed-citation></ref>
<ref id="CIT040"><label>[CLA19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Castillo</surname></string-name>, <string-name><given-names>Jorge</given-names> <surname>L&#x00F3;pez-Moreno</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Aliaga</surname></string-name></person-group>. <article-title>&#x201C;Recent Advances in Fabric Appearance Reproduction&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computers &#x0026; Graphics</italic></source> <volume>84</volume> (<year>2019</year>), <comment>pp.</comment> <fpage>103</fpage>&#x2013;<lpage>121</lpage> (<comment>cit. on pp. 14, 17, 40, 256, 259</comment>).</mixed-citation></ref>
<ref id="CIT041"><label>[CHB21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>Chambon</surname></string-name>, <string-name><given-names>Eric</given-names> <surname>Heitz</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Laurent</given-names> <surname>Belcour</surname></string-name></person-group>. <article-title>&#x201C;Passing Multi-Channel Material Textures to a 3-Channel Loss&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM SIGGRAPH 2021 Talks</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>2</lpage> (<comment>cit. on pp. 64, 66, 67, 128</comment>).</mixed-citation></ref>
<ref id="CIT042"><label>[Cha+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Caroline</given-names> <surname>Chan</surname></string-name>, <string-name><given-names>Shiry</given-names> <surname>Ginosar</surname></string-name>, <string-name><given-names>Tinghui</given-names> <surname>Zhou</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Alexei A.</given-names> <surname>Efros</surname></string-name></person-group>. <article-title>&#x201C;Everybody Dance Now&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <month>Oct</month>. <year>2019</year> (<comment>cit. on p. 85</comment>).</mixed-citation></ref>
<ref id="CIT043"><label>[Che+20a]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Chengqian</given-names> <surname>Che</surname></string-name>, <string-name><given-names>Fujun</given-names> <surname>Luan</surname></string-name>, <string-name><given-names>Shuang</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Kavita</given-names> <surname>Bala</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ioannis</given-names> <surname>Gkioulekas</surname></string-name></person-group>. <article-title>&#x201C;Towards Learning-based Inverse Subsurface Scattering&#x201D;</article-title>. <comment>In</comment>: <source><italic>2020 IEEE International Conference on Computational Photography (ICCP)</italic></source>. <publisher-name>IEEE</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>12</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT044"><label>[Che+22]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Anpei</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Zexiang</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Andreas</given-names> <surname>Geiger</surname></string-name>, <string-name><given-names>Jingyi</given-names> <surname>Yu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hao</given-names> <surname>Su</surname></string-name></person-group>. <chapter-title>&#x201C;Tensorf: Tensorial Radiance Fields&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2022</year>, <comment>pp.</comment> <fpage>333</fpage>&#x2013;<lpage>350</lpage> (<comment>cit. on pp. 203, 212</comment>).</mixed-citation></ref>
<ref id="CIT045"><label>[Che+17a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dongdong</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>Lu</given-names> <surname>Yuan</surname></string-name>, <string-name><given-names>Nenghai</given-names> <surname>Yu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Gang</given-names> <surname>Hua</surname></string-name></person-group>. <article-title>&#x201C;Coherent Online Video Style Transfer&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>1105</fpage>&#x2013;<lpage>1114</lpage> (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT046"><label>[Che+17b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dongdong</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Lu</given-names> <surname>Yuan</surname></string-name>, <string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>Nenghai</given-names> <surname>Yu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Gang</given-names> <surname>Hua</surname></string-name></person-group>. <article-title>&#x201C;Stylebank: An Explicit Representation for Neural Image Style Transfer&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>1897</fpage>&#x2013;<lpage>1906</lpage> (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT047"><label>[Che+20b]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Ting</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Simon</given-names> <surname>Kornblith</surname></string-name>, <string-name><given-names>Mohammad</given-names> <surname>Norouzi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Geoffrey</given-names> <surname>Hinton</surname></string-name></person-group>. <chapter-title>&#x201C;A Simple Framework for Contrastive Learning of Visual Representations&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on <target target-type="page" id="pges_224"/>Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>1597</fpage>&#x2013;<lpage>1607</lpage> (<comment>cit. on pp. 37, 45, 159</comment>).</mixed-citation></ref>
<ref id="CIT048"><label>[CXH21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xinlei</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Saining</given-names> <surname>Xie</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Kaiming</given-names> <surname>He</surname></string-name></person-group>. <article-title>&#x201C;An Empirical Study of Training Self-supervised Vision Transformers&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>9640</fpage>&#x2013;<lpage>9649</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT049"><label>[Che+18]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Yuanqin</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Qian</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Yaping</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Bo</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Meiyun</given-names> <surname>Wang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yusong</given-names> <surname>Lin</surname></string-name></person-group>. <chapter-title>&#x201C;Fine-tuning ResNet for Breast Cancer Classification from mammography&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>The International Conference on Healthcare Science and Engineering</italic></source>. <publisher-name>Springer</publisher-name>. <year>2018</year>, <comment>pp.</comment> <fpage>83</fpage>&#x2013;<lpage>96</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT050"><label>[CNN22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhe</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Shohei</given-names> <surname>Nobuhara</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ko</given-names> <surname>Nishino</surname></string-name></person-group>. <article-title>&#x201C;Invertible Neural BRDF for Object Inverse Rendering&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Pattern Analysis and Machine Intelligence</italic></source> <volume>44</volume>.<issue>12</issue> (<year>2022</year>), <comment>pp.</comment> <fpage>9380</fpage>&#x2013;<lpage>9395</lpage> (<comment>cit. on pp. 186, 203</comment>).</mixed-citation></ref>
<ref id="CIT051"><label>[Cho+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yunjey</given-names> <surname>Choi</surname></string-name>, <string-name><given-names>Minje</given-names> <surname>Choi</surname></string-name>, <string-name><given-names>Munyoung</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>Jung-Woo</given-names> <surname>Ha</surname></string-name>, <string-name><given-names>Sunghun</given-names> <surname>Kim</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jaegul</given-names> <surname>Choo</surname></string-name></person-group>. <article-title>&#x201C;Stargan: Unified Generative Adversarial Networks for Multi-domain Image-to-image Translation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2018</year>, <comment>pp.</comment> <fpage>8789</fpage>&#x2013;<lpage>8797</lpage> (<comment>cit. on p. 77</comment>).</mixed-citation></ref>
<ref id="CIT052"><label>[Cla+90]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Timothy G</given-names> <surname>Clapp</surname></string-name>, <string-name><given-names>Hong</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>Tushar K</given-names> <surname>Ghosh</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jeffrey W</given-names> <surname>Eischen</surname></string-name></person-group>. <article-title>&#x201C;Indirect Measurement of the Moment-curvature Relationship for Fabrics&#x201D;</article-title>. <comment>In</comment>: <source><italic>Textile Research Journal</italic></source> <volume>60</volume>.<issue>9</issue> (<year>1990</year>), <comment>pp.</comment> <fpage>525</fpage>&#x2013;<lpage>533</lpage> (<comment>cit. on pp. 156, 158</comment>).</mixed-citation></ref>
<ref id="CIT053"><label>[Cla+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Petrik</given-names> <surname>Clarberg</surname></string-name>, <string-name><given-names>Wojciech</given-names> <surname>Jarosz</surname></string-name>, <string-name><given-names>Tomas</given-names> <surname>Akenine-M&#x00F6;ller</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Henrik</given-names> <surname>Wann Jensen</surname></string-name></person-group>. <article-title>&#x201C;Wavelet Importance Sampling: Efficiently Evaluating Products of Complex Functions&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (Proceedings of SIGGRAPH)</italic></source> <volume>24</volume>.<issue>3</issue> (<year>2021</year>), <comment>pp.</comment> <fpage>3520</fpage>&#x2013;<lpage>3528</lpage> (<comment>cit. on pp. 185, 189, 195, 199, 200</comment>).</mixed-citation></ref>
<ref id="CIT054"><label>[CTT17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>David</given-names> <surname>Clyde</surname></string-name>, <string-name><given-names>Joseph</given-names> <surname>Teran</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Rasmus</given-names> <surname>Tamstorf</surname></string-name></person-group>. <article-title>&#x201C;Modeling and Data-driven Parameter Estimation for Woven Fabrics&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT055"><label>[Coh+03]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Michael F</given-names> <surname>Cohen</surname></string-name>, <string-name><given-names>Jonathan</given-names> <surname>Shade</surname></string-name>, <string-name><given-names>Stefan</given-names> <surname>Hiller</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Oliver</given-names> <surname>Deussen</surname></string-name></person-group>. <article-title>&#x201C;Wang Tiles for Image and Texture Generation&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>22</volume>.<issue>3</issue> (<year>2003</year>), <comment>pp.</comment> <fpage>287</fpage>&#x2013;<lpage>294</lpage> (<comment>cit. on p. 72</comment>).</mixed-citation></ref>
<ref id="CIT056"><label>[CT82]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Robert L</given-names> <surname>Cook</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Kenneth E.</given-names> <surname>Torrance</surname></string-name></person-group>. <article-title>&#x201C;A Reflectance Model for Computer Graphics&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>1</volume>.<issue>1</issue> (<year>1982</year>), <comment>pp.</comment> <fpage>7</fpage>&#x2013;<lpage>24</lpage> (<comment>cit. on p. 48</comment>).</mixed-citation></ref>
<ref id="CIT057"><label>[Dan01]</label> <mixed-citation publication-type="book"><target target-type="page" id="pges_225"/><person-group person-group-type="author"><string-name><given-names>Kristin J</given-names> <surname>Dana</surname></string-name></person-group>. <chapter-title>&#x201C;BRDF/BTF Measurement Device&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <volume>Vol. 2</volume>. <publisher-name>IEEE</publisher-name>. <year>2001</year>, <comment>pp.</comment> <fpage>460</fpage>&#x2013;<lpage>466</lpage> (<comment>cit. on p. 103</comment>).</mixed-citation></ref>
<ref id="CIT058"><label>[Dan+99]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Kristin J</given-names> <surname>Dana</surname></string-name>, <string-name><given-names>Bram</given-names> <surname>Van Ginneken</surname></string-name>, <string-name><given-names>Shree K</given-names> <surname>Nayar</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan J</given-names> <surname>Koenderink</surname></string-name></person-group>. <article-title>&#x201C;Reflectance and Texture of Real-World Surfaces&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions On Graphics (TOG)</italic></source> <volume>18</volume>.<issue>1</issue> (<year>1999</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>34</lpage> (<comment>cit. on p. 33</comment>).</mixed-citation></ref>
<ref id="CIT059"><label>[Dar+12]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Soheil</given-names> <surname>Darabi</surname></string-name>, <string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name>, <string-name><given-names>Connelly</given-names> <surname>Barnes</surname></string-name>, <string-name><given-names>Dan B</given-names> <surname>Goldman</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Pradeep</given-names> <surname>Sen</surname></string-name></person-group>. <article-title>&#x201C;Image Melding: Combining Inconsistent Images Using Patch-Based Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>31</volume>.<issue>4</issue> (<year>2012</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>10</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT060"><label>[Dav+15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Abe</given-names> <surname>Davis</surname></string-name>, <string-name><given-names>Katherine L</given-names> <surname>Bouman</surname></string-name>, <string-name><given-names>Justin G</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Rubinstein</surname></string-name>, <string-name><given-names>Fredo</given-names> <surname>Durand</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>William T</given-names> <surname>Freeman</surname></string-name></person-group>. <article-title>&#x201C;Visual Vibrometry: Estimating Material Properties from Small Motion in Video&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2015</year>, <comment>pp.</comment> <fpage>5335</fpage>&#x2013;<lpage>5343</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT061"><label>[De 97]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jeremy S</given-names> <surname>De Bonet</surname></string-name></person-group>. <article-title>&#x201C;Multiresolution Sampling Procedure for Analysis and Synthesis of Texture Images&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 24th Annual cCnference on Computer Graphics and Interactive Techniques</italic></source>. <year>1997</year>, <comment>pp.</comment> <fpage>361</fpage>&#x2013;<lpage>368</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT062"><label>[Dek+15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tali</given-names> <surname>Dekel</surname></string-name>, <string-name><given-names>Tomer</given-names> <surname>Michaeli</surname></string-name>, <string-name><given-names>Michal</given-names> <surname>Irani</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>William T.</given-names> <surname>Freeman</surname></string-name></person-group>. <article-title>&#x201C;Revealing and Modifying Non-Local Variations in a Single Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> (<year>2015</year>) (<comment>cit. on p. 93</comment>).</mixed-citation></ref>
<ref id="CIT063"><label>[DH19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>Deliot</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Eric</given-names> <surname>Heitz</surname></string-name></person-group>. <article-title>&#x201C;Procedural Stochastic Textures by Tiling and Blending&#x201D;</article-title>. <comment>In</comment>: <source><italic>GPU Zen</italic></source> <volume>2</volume> (<year>2019</year>) (<comment>cit. on pp. 72, 74, 89&#x2013;91, 95, 99</comment>).</mixed-citation></ref>
<ref id="CIT064"><label>[Den+09]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jia</given-names> <surname>Deng</surname></string-name>, <string-name><given-names>Wei</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Richard</given-names> <surname>Socher</surname></string-name>, <string-name><given-names>Li-Jia</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Kai</given-names> <surname>Li</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Li</given-names> <surname>Fei-Fei</surname></string-name></person-group>. <article-title>&#x201C;Imagenet: A Large-scale Hierarchical Image Database&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</italic></source>. <year>2009</year>, <comment>pp.</comment> <fpage>248</fpage>&#x2013;<lpage>255</lpage> (<comment>cit. on pp. 32, 164, 166</comment>).</mixed-citation></ref>
<ref id="CIT065"><label>[Des+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name>, <string-name><given-names>Miika</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>Fredo</given-names> <surname>Durand</surname></string-name>, <string-name><given-names>George</given-names> <surname>Drettakis</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Adrien</given-names> <surname>Bousseau</surname></string-name></person-group>. <article-title>&#x201C;Single-image SVBRDF Capture with a Rendering-Aware Deep Network&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>37</volume>.<issue>4</issue> (<year>2018</year>) (<comment>cit. on pp. 17, 34, 39, 54, 77, 102, 123, 125, 127, 258</comment>).</mixed-citation></ref>
<ref id="CIT066"><label>[Des+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name>, <string-name><given-names>Miika</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>Fr&#x00E9;do</given-names> <surname>Durand</surname></string-name>, <string-name><given-names>George</given-names> <surname>Drettakis</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Adrien</given-names> <surname>Bousseau</surname></string-name></person-group>. <article-title>&#x201C;Flexible SVBRDF Capture <target target-type="page" id="pges_226"/>with a Multi-Image Deep Network&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 38</volume>. <issue>4</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>13</lpage> (<comment>cit. on pp. 34, 42, 72, 76, 102, 125, 141</comment>).</mixed-citation></ref>
<ref id="CIT067"><label>[DDB20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name>, <string-name><given-names>George</given-names> <surname>Drettakis</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Adrien</given-names> <surname>Bousseau</surname></string-name></person-group>. <article-title>&#x201C;Guided Fine-Tuning for Large-Scale Material Transfer&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 39</volume>. <issue>4</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>91</fpage>&#x2013;<lpage>105</lpage> (<comment>cit. on pp. 35, 38&#x2013;40, 47, 49, 50, 63, 72, 76, 105, 125, 142</comment>).</mixed-citation></ref>
<ref id="CIT068"><label>[DLG21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name>, <string-name><given-names>Yiming</given-names> <surname>Lin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Abhijeet</given-names> <surname>Ghosh</surname></string-name></person-group>. <article-title>&#x201C;Deep Polarization Imaging for 3D Shape and SVBRDF Acquisition&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>15567</fpage>&#x2013;<lpage>15576</lpage> (<comment>cit. on pp. 64, 67, 69, 127, 141</comment>).</mixed-citation></ref>
<ref id="CIT069"><label>[Dia+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Foivos I</given-names> <surname>Diakogiannis</surname></string-name>, <string-name><given-names>Fran&#x00E7;ois</given-names> <surname>Waldner</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Caccetta</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Chen</given-names> <surname>Wu</surname></string-name></person-group>. <article-title>&#x201C;ResUNet-a: A Deep Learning Framework for Semantic Segmentation of Remotely Sensed Data&#x201D;</article-title>. <comment>In</comment>: <source><italic>ISPRS Journal of Photogrammetry and Remote Sensing</italic></source> <volume>162</volume> (<year>2020</year>), <comment>pp.</comment> <fpage>94</fpage>&#x2013;<lpage>114</lpage> (<comment>cit. on pp. 64, 66, 69, 108, 127, 141</comment>).</mixed-citation></ref>
<ref id="CIT070"><label>[Dia+15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Olga</given-names> <surname>Diamanti</surname></string-name>, <string-name><given-names>Connelly</given-names> <surname>Barnes</surname></string-name>, <string-name><given-names>Sylvain</given-names> <surname>Paris</surname></string-name>, <string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Olga</given-names> <surname>Sorkine-Hornung</surname></string-name></person-group>. <article-title>&#x201C;Synthesis of Complex Image Appearance from Limited Exemplars&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>34</volume>.<issue>2</issue> (<year>2015</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>14</lpage> (<comment>cit. on pp. 34, 104</comment>).</mixed-citation></ref>
<ref id="CIT071"><label>[Din+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Keyan</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>Kede</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>Shiqi</given-names> <surname>Wang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Eero P</given-names> <surname>Simoncelli</surname></string-name></person-group>. <article-title>&#x201C;Image Quality Assessment: Unifying Structure and Texture Similarity&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Pattern Analysis and Machine Intelligence</italic></source> (<year>2020</year>) (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT072"><label>[DKB14]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Laurent</given-names> <surname>Dinh</surname></string-name>, <string-name><given-names>David</given-names> <surname>Krueger</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yoshua</given-names> <surname>Bengio</surname></string-name></person-group>. <article-title>&#x201C;Nice: Non-linear Independent Components Estimation&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1410.8516</italic></source> (<year>2014</year>) (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT073"><label>[DSB16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Laurent</given-names> <surname>Dinh</surname></string-name>, <string-name><given-names>Jascha</given-names> <surname>Sohl-Dickstein</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Samy</given-names> <surname>Bengio</surname></string-name></person-group>. <article-title>&#x201C;Density Estimation Using Real NVP&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1605.08803</italic></source> (<year>2016</year>) (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT074"><label>[Dod+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ana</given-names> <surname>Dodik</surname></string-name>, <string-name><given-names>Silvia</given-names> <surname>Sell&#x00E1;n</surname></string-name>, <string-name><given-names>Theodore</given-names> <surname>Kim</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Amanda</given-names> <surname>Phillips</surname></string-name></person-group>. <article-title>&#x201C;Sex and Gender in the Computer Graphics Research Literature&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2206.00480</italic></source> (<year>2022</year>) (<comment>cit. on pp. 178, 179, 182</comment>).</mixed-citation></ref>
<ref id="CIT075"><label>[DZP07]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Weiming</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Ning</given-names> <surname>Zhou</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jean-Claude</given-names> <surname>Paul</surname></string-name></person-group>. <article-title>&#x201C;Optimized Tile-Based Texture Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of Graphics Interface</italic></source> <volume>2007</volume>. <year>2007</year>, <comment>pp.</comment> <fpage>249</fpage>&#x2013;<lpage>256</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT076"><label>[Don19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yue</given-names> <surname>Dong</surname></string-name></person-group>. <article-title>&#x201C;Deep Appearance Modeling: A Survey&#x201D;</article-title>. <comment>In</comment>: <source><italic>Visual Informatics</italic></source> <volume>3</volume>.<issue>2</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>59</fpage>&#x2013;<lpage>68</lpage> (<comment>cit. on pp. 34, 53</comment>).</mixed-citation></ref>
<ref id="CIT077"><label><target target-type="page" id="pges_227"/>[Don+10]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yue</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Jiaping</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Xin</given-names> <surname>Tong</surname></string-name>, <string-name><given-names>John</given-names> <surname>Snyder</surname></string-name>, <string-name><given-names>Yanxiang</given-names> <surname>Lan</surname></string-name>, <string-name><given-names>Moshe</given-names> <surname>Ben-Ezra</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Baining</given-names> <surname>Guo</surname></string-name></person-group>. <article-title>&#x201C;Manifold Bootstrapping for SVBRDF Capture&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>29</volume>.<issue>4</issue> (<year>2010</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>10</lpage> (<comment>cit. on p. 34</comment>).</mixed-citation></ref>
<ref id="CIT078"><label>[DB16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexey</given-names> <surname>Dosovitskiy</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>Brox</surname></string-name></person-group>. <article-title>&#x201C;Generating images with Perceptual Similarity Metrics based on Deep Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2016</year>, <comment>pp.</comment> <fpage>658</fpage>&#x2013;<lpage>666</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT079"><label>[DTD21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Emilien</given-names> <surname>Dupont</surname></string-name>, <string-name><given-names>Yee Whye</given-names> <surname>Teh</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Arnaud</given-names> <surname>Doucet</surname></string-name></person-group>. <article-title>&#x201C;Generative Models as Distributions of Functions&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2102.04776</italic></source> (<year>2021</year>) (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT080"><label>[Dur+19a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Conor</given-names> <surname>Durkan</surname></string-name>, <string-name><given-names>Artur</given-names> <surname>Bekasov</surname></string-name>, <string-name><given-names>Iain</given-names> <surname>Murray</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>George</given-names> <surname>Papamakarios</surname></string-name></person-group>. <article-title>&#x201C;Neural Spline Flows&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <volume>Vol. 32</volume>. <year>2019</year> (<comment>cit. on pp. 185, 191, 192, 194, 196, 197</comment>).</mixed-citation></ref>
<ref id="CIT081"><label>[Dur+19b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Conor</given-names> <surname>Durkan</surname></string-name>, <string-name><given-names>Artur</given-names> <surname>Bekasovs</surname></string-name>, <string-name><given-names>Iain</given-names> <surname>Murray</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Georgios</given-names> <surname>Papamakarios</surname></string-name></person-group>. <article-title>&#x201C;Cubic-Spline Flows&#x201D;</article-title>. <comment>In</comment>: <source><italic>Workshop on Invertible Neural Nets and Normalizing Flows: ICML</italic></source> <volume>2019</volume>. <year>2019</year> (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT082"><label>[Dye+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Joanne</given-names> <surname>Dyer</surname></string-name>, <string-name><given-names>Diego</given-names> <surname>Tamburini</surname></string-name>, <string-name><given-names>Elisabeth R</given-names> <surname>O&#x2019;Connell</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Anna</given-names> <surname>Harrison</surname></string-name></person-group>. <article-title>&#x201C;A Multispectral Imaging Approach Integrated into the Study of Late Antique textiles from Egypt&#x201D;</article-title>. <comment>In</comment>: <source><italic>PLoS One</italic></source> <volume>13</volume>.<issue>10</issue> (<year>2018</year>), <fpage>e0204699</fpage> (<comment>cit. on p. 213</comment>).</mixed-citation></ref>
<ref id="CIT083"><label>[EF01]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexei A</given-names> <surname>Efros</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>William T</given-names> <surname>Freeman</surname></string-name></person-group>. <article-title>&#x201C;Image Quilting for Texture Synthesis and Transfer&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 28th annual conference on Computer Graphics and Interactive Techniques</italic></source>. <year>2001</year>, <comment>pp.</comment> <fpage>341</fpage>&#x2013;<lpage>346</lpage> (<comment>cit. on pp. 72, 74</comment>).</mixed-citation></ref>
<ref id="CIT084"><label>[EL99]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Alexei A</given-names> <surname>Efros</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Thomas K</given-names> <surname>Leung</surname></string-name></person-group>. <chapter-title>&#x201C;Texture Synthesis by Non-parametric Sampling&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <volume>Vol. 2</volume>. <publisher-name>IEEE</publisher-name>. <year>1999</year>, <comment>pp.</comment> <fpage>1033</fpage>&#x2013;<lpage>1038</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT085"><label>[EM17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Michael</given-names> <surname>Elad</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Peyman</given-names> <surname>Milanfar</surname></string-name></person-group>. <article-title>&#x201C;Style Transfer via Texture Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Image Processing</italic></source> <volume>26</volume>.<issue>5</issue> (<year>2017</year>), <comment>pp.</comment> <fpage>2338</fpage>&#x2013;<lpage>2351</lpage> (<comment>cit. on pp. 34, 104</comment>).</mixed-citation></ref>
<ref id="CIT086"><label>[EUD18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Stefan</given-names> <surname>Elfwing</surname></string-name>, <string-name><given-names>Eiji</given-names> <surname>Uchibe</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Kenji</given-names> <surname>Doya</surname></string-name></person-group>. <article-title>&#x201C;Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>Neural Networks</italic></source> <volume>107</volume> (<year>2018</year>), <comment>pp.</comment> <fpage>3</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on pp. 141, 144</comment>).</mixed-citation></ref>
<ref id="CIT087"><label>[End+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yuki</given-names> <surname>Endo</surname></string-name>, <string-name><given-names>Satoshi</given-names> <surname>Iizuka</surname></string-name>, <string-name><given-names>Yoshihiro</given-names> <surname>Kanamori</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jun</given-names> <surname>Mitani</surname></string-name></person-group>. <article-title>&#x201C;Deepprop: Extracting Deep Features from a Single Image for <target target-type="page" id="pges_228"/>Edit Propagation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 35</volume>. <issue>2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2016</year>, <comment>pp.</comment> <fpage>189</fpage>&#x2013;<lpage>201</lpage> (<comment>cit. on p. 33</comment>).</mixed-citation></ref>
<ref id="CIT088"><label>[Fan+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jiahui</given-names> <surname>Fan</surname></string-name>, <string-name><given-names>Beibei</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name>, <string-name><given-names>Jian</given-names> <surname>Yang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ling-Qi</given-names> <surname>Yan</surname></string-name></person-group>. <article-title>&#x201C;Neural Layered BRDFs&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of SIGGRAPH 2022</italic></source>. <year>2022</year> (<comment>cit. on pp. 116, 187</comment>).</mixed-citation></ref>
<ref id="CIT089"><label>[Fen+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xudong</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>Wenchao</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Weiwei</given-names> <surname>Xu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Huamin</given-names> <surname>Wang</surname></string-name></person-group>. <article-title>&#x201C;Learning-Based Bending Stiffness Parameter Estimation by a Drape Tester&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>41</volume>.<issue>6</issue> (<year>2022</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>16</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT090"><label>[Fil+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Filip</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kolafov&#x00E1;</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Havl&#x0131;&#x010D;ek</surname></string-name>, <string-name><given-names>R.</given-names> <surname>V&#x00E1;vra</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Haindl</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><surname>Rushmeier</surname> <given-names>H.</given-names></string-name></person-group> <article-title>&#x201C;Evaluating Physical and Rendered Material Appearance&#x201D;</article-title>. <comment>In</comment>: <source><italic>The Visual Computer (Computer Graphics International 2018)</italic></source> (<year>2018</year>) (<comment>cit. on p. 109</comment>).</mixed-citation></ref>
<ref id="CIT091"><label>[FH08]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ji&#x0159;&#x0131;</given-names> <surname>Filip</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Michal</given-names> <surname>Haindl</surname></string-name></person-group>. <article-title>&#x201C;Bidirectional Texture Function Modeling: A State of the Art Survey&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Pattern Analysis and Machine Intelligence</italic></source> <volume>31</volume>.<issue>11</issue> (<year>2008</year>), <comment>pp.</comment> <fpage>1921</fpage>&#x2013;<lpage>1940</lpage> (<comment>cit. on p. 103</comment>).</mixed-citation></ref>
<ref id="CIT092"><label>[FR22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Michael</given-names> <surname>Fischer</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Tobias</given-names> <surname>Ritschel</surname></string-name></person-group>. <article-title>&#x201C;Plateau-Reduced Differentiable Path Tracing&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2211.17263</italic></source> (<year>2022</year>) (<comment>cit. on p. 202</comment>).</mixed-citation></ref>
<ref id="CIT093"><label>[Fi&#x0161;+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jakub</given-names> <surname>Fi&#x0161;er</surname></string-name>, <string-name><given-names>Ond&#x0159;ej</given-names> <surname>Jamri&#x0161;ka</surname></string-name>, <string-name><given-names>Michal</given-names> <surname>Luk&#x00E1;&#x010D;</surname></string-name>, <string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name>, <string-name><given-names>Paul</given-names> <surname>Asente</surname></string-name>, <string-name><given-names>Jingwan</given-names> <surname>Lu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Sy&#x0060;kora</surname></string-name></person-group>. <article-title>&#x201C;StyLit: Illumination-Guided Example-Based Stylization of 3D Renderings&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>35</volume>.<issue>4</issue> (<year>2016</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on p. 33</comment>).</mixed-citation></ref>
<ref id="CIT094"><label>[Fri+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sara</given-names> <surname>Fridovich-Keil</surname></string-name>, <string-name><given-names>Alex</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>Matthew</given-names> <surname>Tancik</surname></string-name>, <string-name><given-names>Qinhong</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Benjamin</given-names> <surname>Recht</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Angjoo</given-names> <surname>Kanazawa</surname>.</string-name></person-group> <article-title>&#x201C;Plenoxels: Radiance Fields Without Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>5501</fpage>&#x2013;<lpage>5510</lpage> (<comment>cit. on pp. 187, 212</comment>).</mixed-citation></ref>
<ref id="CIT095"><label>[FAW19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Anna</given-names> <surname>Fr&#x00FC;hst&#x00FC;ck</surname></string-name>, <string-name><given-names>Ibraheem</given-names> <surname>Alhashim</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Peter</given-names> <surname>Wonka</surname></string-name></person-group>. <article-title>&#x201C;TileGAN: Synthesis of Large-Scale Non-Homogeneous Textures&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>38</volume>.<issue>4</issue> (<month>Apr</month>. <year>2019</year>) (<comment>cit. on pp. 34, 72, 74, 79&#x2013;81, 104</comment>).</mixed-citation></ref>
<ref id="CIT096"><label>[Fu+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ruigang</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>Qingyong</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Xiaohu</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Yulan</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Yinghui</given-names> <surname>Gao</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Biao</given-names> <surname>Li</surname></string-name></person-group>. <article-title>&#x201C;Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs&#x201D;</article-title>. <comment>In</comment>: <source><italic>31st British Machine Vision Conference 2020, BMVC 2020</italic></source>, <string-name><given-names>BMVA</given-names> <surname>Press</surname></string-name>, <year>2020</year> (<comment>cit. on pp. 167, 168</comment>).</mixed-citation></ref>
<ref id="CIT097"><label>[GG16]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Yarin</given-names> <surname>Gal</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Zoubin</given-names> <surname>Ghahramani</surname></string-name></person-group>. <chapter-title>&#x201C;Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep <target target-type="page" id="pges_229"/>Learning&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2016</year>, <comment>pp.</comment> <fpage>1050</fpage>&#x2013;<lpage>1059</lpage> (<comment>cit. on pp. 123, 126, 130, 141</comment>).</mixed-citation></ref>
<ref id="CIT098"><label>[Gal+12]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Bruno</given-names> <surname>Galerne</surname></string-name>, <string-name><given-names>Ares</given-names> <surname>Lagae</surname></string-name>, <string-name><given-names>Sylvain</given-names> <surname>Lefebvre</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>George</given-names> <surname>Drettakis</surname></string-name></person-group>. <article-title>&#x201C;Gabor Noise by Example&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>31</volume>.<issue>4</issue> (<year>2012</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>9</lpage> (<comment>cit. on p. 72</comment>).</mixed-citation></ref>
<ref id="CIT099"><label>[Gao+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Chen</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>Ayush</given-names> <surname>Saraf</surname></string-name>, <string-name><given-names>Johannes</given-names> <surname>Kopf</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jia-Bin</given-names> <surname>Huang</surname></string-name></person-group>. <article-title>&#x201C;Dynamic View Synthesis From Dynamic Monocular Video&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>5712</fpage>&#x2013;<lpage>5721</lpage> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT100"><label>[Gao+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Duan</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>Xiao</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Yue</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Pieter</given-names> <surname>Peers</surname></string-name>, <string-name><given-names>Kun</given-names> <surname>Xu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Xin</given-names> <surname>Tong</surname></string-name></person-group>. <article-title>&#x201C;Deep Inverse Rendering for High-Resolution SVBRDF Estimation from an Arbitrary Number of Images&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>38</volume>.<issue>4</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>15</lpage> (<comment>cit. on pp. 49, 59, 61, 125, 137, 138, 152, 153</comment>).</mixed-citation></ref>
<ref id="CIT101"><label>[GMX22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Duan</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>Haoyuan</given-names> <surname>Mu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Kun</given-names> <surname>Xu</surname></string-name></person-group>. <article-title>&#x201C;Neural Global Illumination: Interactive In-direct Illumination Prediction under Dynamic Area Lights&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Visualization and Computer Graphics</italic></source> (<year>2022</year>) (<comment>cit. on pp. 105, 186, 203</comment>).</mixed-citation></ref>
<ref id="CIT102"><label>[Gar+14]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name>, <string-name><given-names>Aseem</given-names> <surname>Agarwala</surname></string-name>, <string-name><given-names>Diego</given-names> <surname>Gutierrez</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Aaron</given-names> <surname>Hertzmann</surname></string-name></person-group>. <article-title>&#x201C;A Similarity Measure for Illustration Style&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>33</volume>.<issue>4</issue> (<year>2014</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>9</lpage> (<comment>cit. on pp. 157, 160, 170, 178, 180</comment>).</mixed-citation></ref>
<ref id="CIT103"><label>[Gar+23]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name>, <string-name><given-names>Victor</given-names> <surname>Arellano</surname></string-name>, <string-name><given-names>Carlos</given-names> <surname>Rodriguez-Pardo</surname></string-name>, <string-name><given-names>David</given-names> <surname>Pascual-Hernandez</surname></string-name>, <string-name><given-names>Sergio</given-names> <surname>Suja</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jorge</given-names> <surname>Lopez-Moreno</surname></string-name></person-group>. <article-title>&#x201C;Towards Material Digitization with a Dual-scale Optical System&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> (<year>2023</year>) (<comment>cit. on pp. 15, 16, 24, 102, 109, 269</comment>).</mixed-citation></ref>
<ref id="CIT104"><label>[Gar+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name>, <string-name><given-names>Carlos</given-names> <surname>Rodriguez-Pardo</surname></string-name>, <string-name><given-names>Dan</given-names> <surname>Casas</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jorge</given-names> <surname>Lopez-Moreno</surname></string-name></person-group>. <article-title>&#x201C;A Survey on Intrinsic Images: Delving Deep into Lambert and Beyond&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Journal of Computer Vision</italic></source> (<year>2022</year>) (<comment>cit. on pp. 3, 4, 6, 7, 12, 22, 23, 127, 128, 141, 268</comment>).</mixed-citation></ref>
<ref id="CIT105"><label>[GES22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>James</given-names> <surname>Gardner</surname></string-name>, <string-name><given-names>Bernhard</given-names> <surname>Egger</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>William</given-names> <surname>Alfred Peter Smith</surname></string-name></person-group>. <article-title>&#x201C;Rotation-Equivariant Conditional Spherical Neural Fields for Learning a Natural Illumination Prior&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2022</year> (<comment>cit. on pp. 19, 185, 186, 193, 196, 199, 202, 203, 260</comment>).</mixed-citation></ref>
<ref id="CIT106"><label>[Gar+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Marc-Andr&#x00E9;</given-names> <surname>Gardner</surname></string-name>, <string-name><given-names>Yannick</given-names> <surname>Hold-Geoffroy</surname></string-name>, <string-name><given-names>Kalyan</given-names> <surname>Sunkavalli</surname></string-name>, <string-name><given-names>Christian</given-names> <surname>Gagn&#x00E9;</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jean-Fran&#x00E7;ois</given-names> <surname>Lalonde</surname></string-name></person-group>. <article-title>&#x201C;Deep Parametric Indoor Lighting Estimation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2019</year>, <fpage>pp.</fpage> <fpage>7175</fpage>&#x2013;<lpage>7183</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT107"><label>[GEB15a]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_230"/><person-group person-group-type="author"><string-name><given-names>Leon</given-names> <surname>Gatys</surname></string-name>, <string-name><given-names>Alexander S</given-names> <surname>Ecker</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Matthias</given-names> <surname>Bethge</surname></string-name></person-group>. <article-title>&#x201C;Texture Synthesis using Convolutional Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2015</year>, <comment>pp.</comment> <fpage>262</fpage>&#x2013;<lpage>270</lpage> (<comment>cit. on pp. 74, 95, 128</comment>).</mixed-citation></ref>
<ref id="CIT108"><label>[GEB15b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Leon A</given-names> <surname>Gatys</surname></string-name>, <string-name><given-names>Alexander S</given-names> <surname>Ecker</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Matthias</given-names> <surname>Bethge</surname></string-name></person-group>. <article-title>&#x201C;A Neural Algorithm of Artistic Style&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1508.06576</italic></source> (<year>2015</year>) (<comment>cit. on pp. 32, 47, 48, 124, 128</comment>).</mixed-citation></ref>
<ref id="CIT109"><label>[GEB16a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Leon A</given-names> <surname>Gatys</surname></string-name>, <string-name><given-names>Alexander S</given-names> <surname>Ecker</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Matthias</given-names> <surname>Bethge</surname></string-name></person-group>. <article-title>&#x201C;Image Style Transfer using Convolutional Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2016</year>, <comment>pp.</comment> <fpage>2414</fpage>&#x2013;<lpage>2423</lpage> (<comment>cit. on pp. 74, 160</comment>).</mixed-citation></ref>
<ref id="CIT110"><label>[Gat+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Leon A</given-names> <surname>Gatys</surname></string-name>, <string-name><given-names>Alexander S</given-names> <surname>Ecker</surname></string-name>, <string-name><given-names>Matthias</given-names> <surname>Bethge</surname></string-name>, <string-name><given-names>Aaron</given-names> <surname>Hertzmann</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name></person-group>. <article-title>&#x201C;Controlling Perceptual Factors in Neural Style Transfer&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>3985</fpage>&#x2013;<lpage>3993</lpage> (<comment>cit. on pp. 32, 74</comment>).</mixed-citation></ref>
<ref id="CIT111"><label>[GEB16b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Leon A.</given-names> <surname>Gatys</surname></string-name>, <string-name><given-names>Alexander S.</given-names> <surname>Ecker</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Matthias</given-names> <surname>Bethge</surname></string-name></person-group>. <article-title>&#x201C;Image Style Transfer Using Convolutional Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <volume>Vol. 2016-Decem</volume>. <month>June</month> <year>2016</year>, <comment>pp.</comment> <fpage>2414</fpage>&#x2013;<lpage>2423</lpage> (<comment>cit. on p. 79</comment>).</mixed-citation></ref>
<ref id="CIT112"><label>[Gau+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alban</given-names> <surname>Gauthier</surname></string-name>, <string-name><given-names>Robin</given-names> <surname>Faury</surname></string-name>, <string-name><given-names>J&#x00E9;r&#x00E9;my</given-names> <surname>Levallois</surname></string-name>, <string-name><given-names>Theo</given-names> <surname>Thonat</surname></string-name>, <string-name><given-names>Jean-Marc</given-names> <surname>Thiery</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Tamy</given-names> <surname>Boubekeur</surname></string-name></person-group>. <article-title>&#x201C;MIPNet: Neural Normal-to-Anisotropic-Roughness MIP mapping&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>41</volume>.<issue>6</issue> (<year>2022</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>12</lpage> (<comment>cit. on p. 105</comment>).</mixed-citation></ref>
<ref id="CIT113"><label>[Gaw+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jakob</given-names> <surname>Gawlikowski</surname></string-name>, <string-name><given-names>Cedrique Rovile</given-names> <surname>Njieutcheu Tassi</surname></string-name>, <string-name><given-names>Mohsin</given-names> <surname>Ali</surname></string-name>, <string-name><given-names>Jongseok</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>Matthias</given-names> <surname>Humt</surname></string-name>, <string-name><given-names>Jianxiang</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>Anna</given-names> <surname>Kruspe</surname></string-name>, <string-name><given-names>Rudolph</given-names> <surname>Triebel</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Jung</surname></string-name>, <string-name><given-names>Ribana</given-names> <surname>Roscher</surname></string-name>,</person-group> <etal>et al.</etal> <article-title>&#x201C;A Survey of Uncertainty in Deep Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2107.03342</italic></source> (<year>2021</year>) (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT114"><label>[Geo+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Iliyan</given-names> <surname>Georgiev</surname></string-name>, <string-name><given-names>Thiago</given-names> <surname>Ize</surname></string-name>, <string-name><given-names>Mike</given-names> <surname>Farnsworth</surname></string-name>, <string-name><given-names>Ram&#x00F3;n</given-names> <surname>Montoya-Vozmediano</surname></string-name>, <string-name><given-names>Alan</given-names> <surname>King</surname></string-name>, <string-name><given-names>Brecht</given-names> <surname>Van Lommel</surname></string-name>, <string-name><given-names>Angel</given-names> <surname>Jimenez</surname></string-name>, <string-name><given-names>Oscar</given-names> <surname>Anson</surname></string-name>, <string-name><given-names>Shinji</given-names> <surname>Ogaki</surname></string-name>, <string-name><given-names>Eric</given-names> <surname>Johnston</surname></string-name>,</person-group> <etal>et al.</etal> <article-title>&#x201C;Arnold: A Brute-force Production Path Tracer&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>37</volume>.<issue>3</issue> (<year>2018</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>12</lpage> (<comment>cit. on p. 48</comment>).</mixed-citation></ref>
<ref id="CIT115"><label>[Gil+14]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Guillaume</given-names> <surname>Gilet</surname></string-name>, <string-name><given-names>Basile</given-names> <surname>Sauvage</surname></string-name>, <string-name><given-names>Kenneth</given-names> <surname>Vanhoey</surname></string-name>, <string-name><given-names>Jean-Michel</given-names> <surname>Dischler</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Djamchid</given-names> <surname>Ghazanfarpour</surname></string-name></person-group>. <article-title>&#x201C;Local Random-phase Noise for Procedural Texturing&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>33</volume>.<issue>6</issue> (<year>2014</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on p. 72</comment>).</mixed-citation></ref>
<ref id="CIT116"><label>[Goo+14a]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_231"/><person-group person-group-type="author"><string-name><given-names>Ian</given-names> <surname>Goodfellow</surname></string-name>, <string-name><given-names>Jean</given-names> <surname>Pouget-Abadie</surname></string-name>, <string-name><given-names>Mehdi</given-names> <surname>Mirza</surname></string-name>, <string-name><given-names>Bing</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>David</given-names> <surname>Warde-Farley</surname></string-name>, <string-name><given-names>Sherjil</given-names> <surname>Ozair</surname></string-name>, <string-name><given-names>Aaron</given-names> <surname>Courville</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yoshua</given-names> <surname>Bengio</surname></string-name></person-group>. <article-title>&#x201C;Generative Adversarial Nets&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2014</year>, <comment>pp.</comment> <fpage>2672</fpage>&#x2013;<lpage>2680</lpage> (<comment>cit. on p. 39</comment>).</mixed-citation></ref>
<ref id="CIT117"><label>[Goo+14b]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Ian J.</given-names> <surname>Goodfellow</surname></string-name>, <string-name><given-names>Jean</given-names> <surname>Pouget-Abadie</surname></string-name>, <string-name><given-names>Mehdi</given-names> <surname>Mirza</surname></string-name>, <string-name><given-names>Bing</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>David</given-names> <surname>Warde-Farley</surname></string-name>, <string-name><given-names>Sherjil</given-names> <surname>Ozair</surname></string-name>, <string-name><given-names>Aaron</given-names> <surname>Courville</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yoshua</given-names> <surname>Bengio</surname></string-name></person-group>. <chapter-title>&#x201C;Generative Adversarial Nets&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <volume>Vol. 3</volume>. <month>January</month>. <publisher-name>Neural information processing systems foundation</publisher-name>, <month>June</month> <year>2014</year>, <comment>pp.</comment> <fpage>2672</fpage>&#x2013;<lpage>2680</lpage> (<comment>cit. on p. 79</comment>).</mixed-citation></ref>
<ref id="CIT118"><label>[Gri+03]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Eitan</given-names> <surname>Grinspun</surname></string-name>, <string-name><given-names>Anil N</given-names> <surname>Hirani</surname></string-name>, <string-name><given-names>Mathieu</given-names> <surname>Desbrun</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Peter</given-names> <surname>Schr&#x00F6;der</surname></string-name></person-group>. <chapter-title>&#x201C;Discrete Shells&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the 2003 ACM SIGGRAPH/Eurographics symposium on Computer animation</italic></source>. <publisher-name>Citeseer</publisher-name>. <year>2003</year>, <comment>pp.</comment> <fpage>62</fpage>&#x2013;<lpage>67</lpage> (<comment>cit. on pp. 161, 162</comment>).</mixed-citation></ref>
<ref id="CIT119"><label>[Gro+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Aditya</given-names> <surname>Grover</surname></string-name>, <string-name><given-names>Christopher</given-names> <surname>Chute</surname></string-name>, <string-name><given-names>Rui</given-names> <surname>Shu</surname></string-name>, <string-name><given-names>Zhangjie</given-names> <surname>Cao</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Stefano</given-names> <surname>Ermon</surname></string-name></person-group>. <article-title>&#x201C;Align-flow: Cycle Consistent Learning from Multiple Domains via Normalizing Flows&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the AAAI Conference on Artificial Intelligence</italic></source>. <volume>Vol. 34</volume>. <issue>04</issue>. <year>2020</year>, <fpage>pp.</fpage> <fpage>4028</fpage>&#x2013;<lpage>4035</lpage> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT120"><label>[Gu+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Shuyang</given-names> <surname>Gu</surname></string-name>, <string-name><given-names>Congliang</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Lu</given-names> <surname>Yuan</surname></string-name></person-group>. <article-title>&#x201C;Arbitrary Style Transfer with Deep Feature Reshuffle&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2018</year>, <comment>pp.</comment> <fpage>8222</fpage>&#x2013;<lpage>8231</lpage> (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT121"><label>[Gua+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Darya</given-names> <surname>Guarnera</surname></string-name>, <string-name><given-names>Giuseppe Claudio</given-names> <surname>Guarnera</surname></string-name>, <string-name><given-names>Abhijeet</given-names> <surname>Ghosh</surname></string-name>, <string-name><given-names>Cornelia</given-names> <surname>Denk</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mashhuda</given-names> <surname>Glencross</surname></string-name></person-group>. <article-title>&#x201C;BRDF Representation and Acquisition&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 35</volume>. <issue>2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2016</year>, <comment>pp.</comment> <fpage>625</fpage>&#x2013;<lpage>650</lpage> (<comment>cit. on p. 34</comment>).</mixed-citation></ref>
<ref id="CIT122"><label>[Gua+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Giuseppe Claudio</given-names> <surname>Guarnera</surname></string-name>, <string-name><given-names>Dar&#x2019;ya</given-names> <surname>Guarnera</surname></string-name>, <string-name><given-names>Gregory J</given-names> <surname>Ward</surname></string-name>, <string-name><given-names>Mashhuda</given-names> <surname>Glencross</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ian</given-names> <surname>Hall</surname></string-name></person-group>. <article-title>&#x201C;BxDF Material Acquisition, Representation, and Rendering for VR and Design&#x201D;</article-title>. <comment>In</comment>: <source><italic>SIGGRAPH Asia 2019 Courses</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>21</lpage> (<comment>cit. on p. 3</comment>).</mixed-citation></ref>
<ref id="CIT123"><label>[Gue+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P</given-names> <surname>Guehl</surname></string-name>, <string-name><given-names>R</given-names> <surname>All&#x00E8;gre</surname></string-name>, <string-name><given-names>J-M</given-names> <surname>Dischler</surname></string-name>, <string-name><given-names>B</given-names> <surname>Benes</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>E</given-names> <surname>Galin</surname></string-name></person-group>. <article-title>&#x201C;Semi-Procedural Textures Using Point Process Texture Basis Functions&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 39</volume>. <issue>4</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>159</fpage>&#x2013;<lpage>171</lpage> (<comment>cit. on pp. 34, 72, 92, 104</comment>).</mixed-citation></ref>
<ref id="CIT124"><label>[Gui+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Geoffrey</given-names> <surname>Guingo</surname></string-name>, <string-name><given-names>Basile</given-names> <surname>Sauvage</surname></string-name>, <string-name><given-names>Jean-Michel</given-names> <surname>Dischler</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Marie-Paule</given-names> <surname>Cani</surname></string-name></person-group>. <article-title>&#x201C;Bilayer Textures: A Model for Synthesis and Deformation of Composite Textures&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 36</volume>. <issue>4</issue>. <year>2017</year>, <comment>pp.</comment> <fpage>111</fpage>&#x2013;<lpage>122</lpage> (<comment>cit. on p. 72</comment>).</mixed-citation></ref>
<ref id="CIT125"><label>[Guo+21]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_232"/><person-group person-group-type="author"><string-name><given-names>Jie</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Shuichang</given-names> <surname>Lai</surname></string-name>, <string-name><given-names>Chengzhi</given-names> <surname>Tao</surname></string-name>, <string-name><given-names>Yuelong</given-names> <surname>Cai</surname></string-name>, <string-name><given-names>Lei</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Yanwen</given-names> <surname>Guo</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ling-Qi</given-names> <surname>Yan</surname></string-name></person-group>. <article-title>&#x201C;Highlight-Aware Two-Stream Network for Single-Image SVBRDF Acquisition&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>40</volume>.<issue>4</issue> (<year>2021</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>14</lpage> (<comment>cit. on pp. 123, 125</comment>).</mixed-citation></ref>
<ref id="CIT126"><label>[Guo+20a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yu</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name>, <string-name><given-names>Lingqi</given-names> <surname>Yan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Shuang</given-names> <surname>Zhao</surname></string-name></person-group>. <article-title>&#x201C;A Bayesian Inference Framework for Procedural Material Parameter Estimation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 39</volume>. <issue>7</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>255</fpage>&#x2013;<lpage>266</lpage> (<comment>cit. on pp. 125, 126</comment>).</mixed-citation></ref>
<ref id="CIT127"><label>[Guo+20b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yu</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Cameron</given-names> <surname>Smith</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name>, <string-name><given-names>Kalyan</given-names> <surname>Sunkavalli</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Shuang</given-names> <surname>Zhao</surname></string-name></person-group>. <article-title>&#x201C;MaterialGAN: Reflectance Capture using a Generative SVBRDF Model&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>39</volume>.<issue>6</issue> (<year>2020</year>), <volume>254</volume>:<fpage>1</fpage>&#x2013;<lpage>254</lpage>:<issue>13</issue> (<comment>cit. on pp. 17, 35, 42, 72, 76, 102, 123, 125, 258</comment>).</mixed-citation></ref>
<ref id="CIT128"><label>[Guo+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yunhui</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Honghui</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>Abhishek</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>Kristen</given-names> <surname>Grauman</surname></string-name>, <string-name><given-names>Tajana</given-names> <surname>Rosing</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Rogerio</given-names> <surname>Feris</surname></string-name></person-group>. <article-title>&#x201C;Spottune: Transfer Learning through Adaptive Fine-tuning&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>4805</fpage>&#x2013;<lpage>4814</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT129"><label>[HDL16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>David</given-names> <surname>Ha</surname></string-name>, <string-name><given-names>Andrew</given-names> <surname>Dai</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Quoc V</given-names> <surname>Le</surname></string-name></person-group>. <article-title>&#x201C;Hypernetworks&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1609.09106</italic></source> (<year>2016</year>) (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT130"><label>[HMV09]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Simon</given-names> <surname>Haegler</surname></string-name>, <string-name><given-names>Pascal</given-names> <surname>M&#x00FC;ller</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Luc</given-names> <surname>Van Gool</surname></string-name></person-group>. <article-title>&#x201C;Procedural Modeling for Digital Cultural Heritage&#x201D;</article-title>. <comment>In</comment>: <source><italic>EURASIP Journal on Image and Video Processing</italic></source> <volume>2009</volume> (<year>2009</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>1</lpage> (<comment>cit. on p. 213</comment>).</mixed-citation></ref>
<ref id="CIT131"><label>[Hai12]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>V&#x00E1;vra</surname> <given-names>R.</given-names></string-name> <string-name><surname>Haindl</surname> <given-names>M.</given-names></string-name> <string-name><surname>Filip</surname> <given-names>J.</given-names></string-name></person-group> <article-title>&#x201C;Digital Material Appearance: the Curse of Tera-Bytes&#x201D;</article-title>. <comment>In</comment>: <source><italic>ERCIM News</italic></source> <volume>90</volume> (<year>2012</year>), <comment>pp.</comment> <fpage>49</fpage>&#x2013;<lpage>50</lpage> (<comment>cit. on pp. 108, 109, 111</comment>).</mixed-citation></ref>
<ref id="CIT132"><label>[Ham+21]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Hendrik</given-names> <surname>Hameeuw</surname></string-name>, <string-name><given-names>Godelieve</given-names> <surname>Watteeuw</surname></string-name>, <string-name><given-names>Bruno</given-names> <surname>Vandermeulen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Marc</given-names> <surname>Proesmans</surname></string-name></person-group>. <article-title>&#x201C;The Painted Panels of the Early Sixteenth Century Mechelen Enclosed Gardens Art Technical Examination with the Photometric Stereo: White Light and Multispectral Microdomes&#x201D;</article-title>. <comment>In</comment>: <source><italic>Papers Presented at the Twentieth Symposium for the Study of Underdrawing and Technology in Painting held in Mechelen and Leuven, 11-13 January 2017</italic></source>. <volume>Vol. 20</volume>. <publisher-name>Peeters Publishers; Leuven</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>165</fpage>&#x2013;<lpage>175</lpage> (<comment>cit. on p. 213</comment>).</mixed-citation></ref>
<ref id="CIT133"><label>[Han+22a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jiyeon</given-names> <surname>Han</surname></string-name>, <string-name><given-names>Hwanil</given-names> <surname>Choi</surname></string-name>, <string-name><given-names>Yunjey</given-names> <surname>Choi</surname></string-name>, <string-name><given-names>Junho</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>Jung-Woo</given-names> <surname>Ha</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jaesik</given-names> <surname>Choi</surname></string-name></person-group>. <article-title>&#x201C;Rarity Score: A New Metric to Evaluate the Uncommonness of Synthesized Images&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2206.08549</italic></source> (<year>2022</year>) (<comment>cit. on p. 128</comment>).</mixed-citation></ref>
<ref id="CIT134"><label>[Han+22b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Kai</given-names> <surname>Han</surname></string-name>, <string-name><given-names>Yunhe</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Hanting</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Xinghao</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Jianyuan</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Zhenhua</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Yehui</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>An</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>Chunjing</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Yixing</given-names> <surname>Xu</surname></string-name></person-group>, <etal>et al.</etal> <target target-type="page" id="pges_233"/><article-title>&#x201C;A Survey on Vision Transformer&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Pattern Analysis and Machine Intelligence</italic></source> (<year>2022</year>) (<comment>cit. on p. 164</comment>).</mixed-citation></ref>
<ref id="CIT135"><label>[Haq+23]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ayaan</given-names> <surname>Haque</surname></string-name>, <string-name><given-names>Matthew</given-names> <surname>Tancik</surname></string-name>, <string-name><given-names>Alexei</given-names> <surname>Efros</surname></string-name>, <string-name><given-names>Aleksander</given-names> <surname>Holynski</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Angjoo</given-names> <surname>Kanazawa</surname></string-name></person-group>. <article-title>&#x201C;Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2303.12789</italic></source> (<year>2023</year>) (<comment>cit. on p. 116</comment>).</mixed-citation></ref>
<ref id="CIT136"><label>[HHM22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jon</given-names> <surname>Hasselgren</surname></string-name>, <string-name><given-names>Nikolai</given-names> <surname>Hofmann</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jacob</given-names> <surname>Munkberg</surname></string-name></person-group>. <article-title>&#x201C;Shape, Light &#x0026; Material Decomposition from Images Using Monte Carlo Rendering and Denoising&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2206.03380</italic></source> (<year>2022</year>) (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT137"><label>[HFM10]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Vlastimil</given-names> <surname>Havran</surname></string-name>, <string-name><given-names>Jir&#x0131;</given-names> <surname>Filip</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Karol</given-names> <surname>Myszkowski</surname></string-name></person-group>. <article-title>&#x201C;Bidirectional Texture Function Compression Based on Multi-level Vector Quantization&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 29</volume>. <issue>1</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2010</year>, <comment>pp.</comment> <fpage>175</fpage>&#x2013;<lpage>190</lpage> (<comment>cit. on p. 103</comment>).</mixed-citation></ref>
<ref id="CIT138"><label>[He+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Kaiming</given-names> <surname>He</surname></string-name>, <string-name><given-names>Xinlei</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Saining</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>Yanghao</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Piotr</given-names> <surname>Doll&#x00E1;r</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ross</given-names> <surname>Girshick</surname></string-name></person-group>. <article-title>&#x201C;Masked Autoencoders are Scalable Vision Learners&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>16000</fpage>&#x2013;<lpage>16009</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT139"><label>[He+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Kaiming</given-names> <surname>He</surname></string-name>, <string-name><given-names>Xiangyu</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Shaoqing</given-names> <surname>Ren</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jian</given-names> <surname>Sun</surname></string-name></person-group>. <article-title>&#x201C;Deep Residual Learning for Image Recognition&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <volume>Vol. 2016-Decem</volume>. <year>2016</year>, <comment>pp.</comment> <fpage>770</fpage>&#x2013;<lpage>778</lpage> (<comment>cit. on pp. 64, 66, 69, 78, 94, 108, 109, 127, 141, 164, 176</comment>).</mixed-citation></ref>
<ref id="CIT140"><label>[He+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Mingming</given-names> <surname>He</surname></string-name>, <string-name><given-names>Dongdong</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>Pedro V</given-names> <surname>Sander</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Lu</given-names> <surname>Yuan</surname></string-name></person-group>. <article-title>&#x201C;Deep Exemplar-Based Colorization&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>37</volume>.<issue>4</issue> (<year>2018</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>16</lpage> (<comment>cit. on p. 33</comment>).</mixed-citation></ref>
<ref id="CIT141"><label>[He+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Mingming</given-names> <surname>He</surname></string-name>, <string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>Dongdong</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Lu</given-names> <surname>Yuan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Pedro V</given-names> <surname>Sander</surname></string-name></person-group>. <article-title>&#x201C;Progressive Color Transfer with Dense Semantic Correspondences&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>38</volume>.<issue>2</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>18</lpage> (<comment>cit. on p. 33</comment>).</mixed-citation></ref>
<ref id="CIT142"><label>[HB95]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>David J</given-names> <surname>Heeger</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>James R</given-names> <surname>Bergen</surname></string-name></person-group>. <article-title>&#x201C;Pyramid-based Texture Analysis/Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques</italic></source>. <year>1995</year>, <comment>pp.</comment> <fpage>229</fpage>&#x2013;<lpage>238</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT143"><label>[Hei20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Eric</given-names> <surname>Heitz</surname></string-name></person-group>. <article-title>&#x201C;Can&#x2019;t Invert the CDF? The Triangle-Cut Parameterization of the Region under the Curve&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source> <volume>39</volume>.<issue>4</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>121</fpage>&#x2013;<lpage>132</lpage> (<comment>cit. on p. 185</comment>).</mixed-citation></ref>
<ref id="CIT144"><label>[HN18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Eric</given-names> <surname>Heitz</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Fabrice</given-names> <surname>Neyret</surname></string-name></person-group>. <article-title>&#x201C;High-Performance By-Example Noise Using a Histogram-Preserving Blending Operator&#x201D;</article-title>. <target target-type="page" id="pges_234"/><comment>In</comment>: <source><italic>Proceedings of the ACM on Computer Graphics and Interactive Techniques</italic></source> <volume>1</volume>.<issue>2</issue> (<year>2018</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>25</lpage> (<comment>cit. on p. 72</comment>).</mixed-citation></ref>
<ref id="CIT145"><label>[Hei+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Eric</given-names> <surname>Heitz</surname></string-name>, <string-name><given-names>Kenneth</given-names> <surname>Vanhoey</surname></string-name>, <string-name><given-names>Thomas</given-names> <surname>Chambon</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Laurent</given-names> <surname>Belcour</surname></string-name></person-group>. <article-title>&#x201C;A Sliced Wasserstein Loss for Neural Texture Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>9412</fpage>&#x2013;<lpage>9420</lpage> (<comment>cit. on pp. 75, 93</comment>).</mixed-citation></ref>
<ref id="CIT146"><label>[Hel+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Leonhard</given-names> <surname>Helminger</surname></string-name>, <string-name><given-names>Abdelaziz</given-names> <surname>Djelouah</surname></string-name>, <string-name><given-names>Markus</given-names> <surname>Gross</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Christopher</given-names> <surname>Schroers</surname></string-name></person-group>. <article-title>&#x201C;Lossy Image Compression with Normalizing Flows&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2008.10486</italic></source> (<year>2020</year>) (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT147"><label>[HG16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dan</given-names> <surname>Hendrycks</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Kevin</given-names> <surname>Gimpel</surname></string-name></person-group>. <article-title>&#x201C;Gaussian Error Linear Units (GELUs)&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1606.08415</italic></source> (<year>2016</year>) (<comment>cit. on pp. 109, 119</comment>).</mixed-citation></ref>
<ref id="CIT148"><label>[Hen+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Philipp</given-names> <surname>Henzler</surname></string-name>, <string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name>, <string-name><given-names>Niloy J</given-names> <surname>Mitra</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Tobias</given-names> <surname>Ritschel</surname></string-name></person-group>. <article-title>&#x201C;Generative Modelling of BRDF Textures from Flash Images&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (Proc. SIGGRAPH Asia)</italic></source> <volume>40</volume>.<issue>6</issue> (<year>2021</year>) (<comment>cit. on pp. 17, 76, 102, 123, 124, 137, 138, 143, 151&#x2013;153, 258</comment>).</mixed-citation></ref>
<ref id="CIT149"><label>[Her+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Amir</given-names> <surname>Hertz</surname></string-name>, <string-name><given-names>Rana</given-names> <surname>Hanocka</surname></string-name>, <string-name><given-names>Raja</given-names> <surname>Giryes</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Cohen-Or</surname></string-name></person-group>. <article-title>&#x201C;Deep Geometric Texture Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>39</volume>.<issue>4</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>108</fpage>&#x2013;<lpage>1</lpage> (<comment>cit. on p. 72</comment>).</mixed-citation></ref>
<ref id="CIT150"><label>[Her+01]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Aaron</given-names> <surname>Hertzmann</surname></string-name>, <string-name><given-names>Charles E</given-names> <surname>Jacobs</surname></string-name>, <string-name><given-names>Nuria</given-names> <surname>Oliver</surname></string-name>, <string-name><given-names>Brian</given-names> <surname>Curless</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>David H</given-names> <surname>Salesin</surname></string-name></person-group>. <article-title>&#x201C;Image Analogies&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 28th annual conference on Computer graphics and interactive techniques</italic></source>. <year>2001</year>, <comment>pp.</comment> <fpage>327</fpage>&#x2013;<lpage>340</lpage> (<comment>cit. on pp. 31, 32, 46, 56, 57, 60</comment>).</mixed-citation></ref>
<ref id="CIT151"><label>[HS05]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Aaron</given-names> <surname>Hertzmann</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Steven M</given-names> <surname>Seitz</surname></string-name></person-group>. <article-title>&#x201C;Example-based Photometric Stereo: Shape Reconstruction with General, Varying BRDFs&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Pattern Analysis and Machine Intelligence</italic></source> <volume>27</volume>.<issue>8</issue> (<year>2005</year>), <comment>pp.</comment> <fpage>1254</fpage>&#x2013;<lpage>1264</lpage> (<comment>cit. on p. 34</comment>).</mixed-citation></ref>
<ref id="CIT152"><label>[HS03]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Aaron</given-names> <surname>Hertzmann</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Steven M</given-names> <surname>Seitz</surname></string-name></person-group>. <article-title>&#x201C;Shape and Materials by example: A Photometric Stereo Approach&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <volume>Vol. 1</volume>. <publisher-name>IEEE</publisher-name>. <year>2003</year>, <fpage>pp.</fpage> <fpage>I</fpage>&#x2013;<lpage>I</lpage> (<comment>cit. on p. 34</comment>).</mixed-citation></ref>
<ref id="CIT153"><label>[Heu+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Martin</given-names> <surname>Heusel</surname></string-name>, <string-name><given-names>Hubert</given-names> <surname>Ramsauer</surname></string-name>, <string-name><given-names>Thomas</given-names> <surname>Unterthiner</surname></string-name>, <string-name><given-names>Bernhard</given-names> <surname>Nessler</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Sepp</given-names> <surname>Hochreiter</surname></string-name></person-group>. <article-title>&#x201C;Gans Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1706.08500</italic></source> (<year>2017</year>) (<comment>cit. on p. 85</comment>).</mixed-citation></ref>
<ref id="CIT154"><label>[Hil+15]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Stephen</given-names> <surname>Hill</surname></string-name>, <string-name><given-names>Stephen</given-names> <surname>McAuley</surname></string-name>, <string-name><given-names>Brent</given-names> <surname>Burley</surname></string-name></person-group>, <etal>et al.</etal> <chapter-title>&#x201C;Physically Based Shading in Theory and Practice&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>ACM SIGGRAPH <target target-type="page" id="pges_235"/>2015 Courses</italic></source>. <comment>SIGGRAPH &#x2019;15</comment>. <publisher-loc>Los Angeles, California</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>, <year>2015</year> (<comment>cit. on p. 11</comment>).</mixed-citation></ref>
<ref id="CIT155"><label>[HS06]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Geoffrey E</given-names> <surname>Hinton</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ruslan R</given-names> <surname>Salakhutdinov</surname></string-name></person-group>. <article-title>&#x201C;Reducing the Dimensionality of Data with Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>science</italic></source> <volume>313</volume>.<issue>5786</issue> (<year>2006</year>), <comment>pp.</comment> <fpage>504</fpage>&#x2013;<lpage>507</lpage> (<comment>cit. on p. 103</comment>).</mixed-citation></ref>
<ref id="CIT156"><label>[Hin+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tobias</given-names> <surname>Hinz</surname></string-name>, <string-name><given-names>Matthew</given-names> <surname>Fisher</surname></string-name>, <string-name><given-names>Oliver</given-names> <surname>Wang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Stefan</given-names> <surname>Wermter</surname></string-name></person-group>. <article-title>&#x201C;Improved Techniques for Training Single-Image GANs&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings - 2021 IEEE Winter Conference on Applications of Computer Vision, WACV</italic></source> <volume>2021</volume>. <month>Jan</month>. <year>2021</year>, <comment>pp.</comment> <fpage>1300</fpage>&#x2013;<lpage>1309</lpage> (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT157"><label>[HJA20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jonathan</given-names> <surname>Ho</surname></string-name>, <string-name><given-names>Ajay</given-names> <surname>Jain</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Pieter</given-names> <surname>Abbeel</surname></string-name></person-group>. <article-title>&#x201C;Denoising Diffusion Probabilistic Models&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>33</volume> (<year>2020</year>), <comment>pp.</comment> <fpage>6840</fpage>&#x2013;<lpage>6851</lpage> (<comment>cit. on pp. 127, 209</comment>).</mixed-citation></ref>
<ref id="CIT158"><label>[HS13]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Shahera</given-names> <surname>Hossain</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Seiichi</given-names> <surname>Serikawa</surname></string-name></person-group>. <article-title>&#x201C;Texture Databases&#x2013;A Comprehensive Survey&#x201D;</article-title>. <comment>In</comment>: <source><italic>Pattern Recognition Letters</italic></source> <volume>34</volume>.<issue>15</issue> (<year>2013</year>), <comment>pp.</comment> <fpage>2007</fpage>&#x2013;<lpage>2022</lpage> (<comment>cit. on p. 109</comment>).</mixed-citation></ref>
<ref id="CIT159"><label>[Hu+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Bingyang</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Jie</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Yanjun</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Mengtian</given-names> <surname>Li</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yanwen</given-names> <surname>Guo</surname></string-name></person-group>. <article-title>&#x201C;DeepBRDF: A Deep Representation for Manipulating Measured BRDF&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 39</volume>. <issue>2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>157</fpage>&#x2013;<lpage>166</lpage> (<comment>cit. on pp. 105, 203</comment>).</mixed-citation></ref>
<ref id="CIT160"><label>[Hu+22a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dongting</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Liuhua</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>Tingjin</given-names> <surname>Chu</surname></string-name>, <string-name><given-names>Xiaoxing</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Yinian</given-names> <surname>Mao</surname></string-name>, <string-name><given-names>Howard</given-names> <surname>Bondell</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mingming</given-names> <surname>Gong</surname></string-name></person-group>. <article-title>&#x201C;Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <year>2022</year> (<comment>cit. on pp. 126, 130</comment>).</mixed-citation></ref>
<ref id="CIT161"><label>[HSS18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jie</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Li</given-names> <surname>Shen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Gang</given-names> <surname>Sun</surname></string-name></person-group>. <article-title>&#x201C;Squeeze-and-excitation Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2018</year>, <fpage>pp.</fpage> <fpage>7132</fpage>&#x2013;<lpage>7141</lpage> (<comment>cit. on p. 164</comment>).</mixed-citation></ref>
<ref id="CIT162"><label>[HXP20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wei</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Lechao</given-names> <surname>Xiao</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jeffrey</given-names> <surname>Pennington</surname></string-name></person-group>. <article-title>&#x201C;Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2001.05992</italic></source> (<year>2020</year>) (<comment>cit. on p. 141</comment>).</mixed-citation></ref>
<ref id="CIT163"><label>[HDR19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yiwei</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Julie</given-names> <surname>Dorsey</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Holly</given-names> <surname>Rushmeier</surname></string-name></person-group>. <article-title>&#x201C;A Novel Framework for Inverse Procedural Texture Modeling&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>38</volume>.<issue>6</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>14</lpage> (<comment>cit. on p. 72</comment>).</mixed-citation></ref>
<ref id="CIT164"><label>[Hu+22b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yiwei</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name>, <string-name><given-names>Paul</given-names> <surname>Guerrero</surname></string-name>, <string-name><given-names>Holly</given-names> <surname>Rushmeier</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name></person-group>. <article-title>&#x201C;Controlling Material Appearance by Examples&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 41</volume>. <issue>4</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2022</year>, <comment>pp.</comment> <fpage>117</fpage>&#x2013;<lpage>128</lpage> (<comment>cit. on p. 105</comment>).</mixed-citation></ref>
<ref id="CIT165"><label>[Hu+22c]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_236"/><person-group person-group-type="author"><string-name><given-names>Yiwei</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Chengan</given-names> <surname>He</surname></string-name>, <string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name>, <string-name><given-names>Julie</given-names> <surname>Dorsey</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Holly</given-names> <surname>Rushmeier</surname></string-name></person-group>. <article-title>&#x201C;An Inverse Procedural Modeling Pipeline for SVBRDF Maps&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>41</volume>.<issue>2</issue> (<year>2022</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>17</lpage> (<comment>cit. on p. 125</comment>).</mixed-citation></ref>
<ref id="CIT166"><label>[Hu+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yuanming</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Luke</given-names> <surname>Anderson</surname></string-name>, <string-name><given-names>Tzu-Mao</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Qi</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>Nathan</given-names> <surname>Carr</surname></string-name>, <string-name><given-names>Jonathan</given-names> <surname>Ragan-Kelley</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Fredo</given-names> <surname>Durand</surname></string-name></person-group>. <article-title>&#x201C;Diff Taichi: Differentiable Programming for Physical Simulation&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2019</year> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT167"><label>[Hua+18a]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Chin-Wei</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>David</given-names> <surname>Krueger</surname></string-name>, <string-name><given-names>Alexandre</given-names> <surname>Lacoste</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Aaron</given-names> <surname>Courville</surname></string-name></person-group>. <chapter-title>&#x201C;Neural Autoregressive Flows&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2018</year>, <comment>pp.</comment> <fpage>2078</fpage>&#x2013;<lpage>2087</lpage> (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT168"><label>[HB17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xun</given-names> <surname>Huang</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Serge</given-names> <surname>Belongie</surname></string-name></person-group>. <article-title>&#x201C;Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>1501</fpage>&#x2013;<lpage>1510</lpage> (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT169"><label>[Hua+18b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xun</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Ming-Yu</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Serge</given-names> <surname>Belongie</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Kautz</surname></string-name></person-group>. <article-title>&#x201C;Multimodal Unsupervised Image-to-image Translation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <month>Sept</month>. <year>2018</year> (<comment>cit. on p. 85</comment>).</mixed-citation></ref>
<ref id="CIT170"><label>[Hua+19]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Zhiyuan</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Mansur</given-names> <surname>Arief</surname></string-name>, <string-name><given-names>Henry</given-names> <surname>Lam</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ding</given-names> <surname>Zhao</surname></string-name></person-group>. <chapter-title>&#x201C;Evaluation Uncertainty in Data-Driven Self-Driving Testing&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>2019 IEEE Intelligent Transportation Systems Conference (ITSC)</italic></source>. <publisher-name>IEEE</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>1902</fpage>&#x2013;<lpage>1907</lpage> (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT171"><label>[HEW17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Markus</given-names> <surname>Huber</surname></string-name>, <string-name><given-names>Bernhard</given-names> <surname>Eberhardt</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Weiskopf</surname></string-name></person-group>. <article-title>&#x201C;Cloth Animation Retrieval Using a Motion-shape Signature&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Computer Graphics and Applications</italic></source> <volume>37</volume>.<issue>6</issue> (<year>2017</year>), <comment>pp.</comment> <fpage>52</fpage>&#x2013;<lpage>64</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT172"><label>[HZ21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Drew A</given-names> <surname>Hudson</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>C Lawrence</given-names> <surname>Zitnick</surname></string-name></person-group>. <article-title>&#x201C;Generative Adversarial Transformers&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2103.01209</italic></source> (<year>2021</year>) (<comment>cit. on p. 93</comment>).</mixed-citation></ref>
<ref id="CIT173"><label>[Hwa+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Inseung</given-names> <surname>Hwang</surname></string-name>, <string-name><given-names>Daniel S</given-names> <surname>Jeon</surname></string-name>, <string-name><given-names>Adolfo</given-names> <surname>Mu&#x00F1;oz</surname></string-name>, <string-name><given-names>Diego</given-names> <surname>Gutierrez</surname></string-name>, <string-name><given-names>Xin</given-names> <surname>Tong</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Min H</given-names> <surname>Kim</surname></string-name></person-group>. <article-title>&#x201C;Sparse Ellipsometry: Portable Acquisition of Polarimetric SVBRDF and Shape with Unstructured Flash Photography&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>41</volume>.<issue>4</issue> (<year>2022</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>14</lpage> (<comment>cit. on p. 129</comment>).</mixed-citation></ref>
<ref id="CIT174"><label>[Ike81]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Katsushi</given-names> <surname>Ikeuchi</surname></string-name></person-group>. <article-title>&#x201C;Determining Surface Orientations of Specular Surfaces by Using the Photometric Stereo Method&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Pattern Analysis and Machine Intelligence</italic></source> <volume>6</volume> (<year>1981</year>), <comment>pp.</comment> <fpage>661</fpage>&#x2013;<lpage>669</lpage> (<comment>cit. on pp. 34, 40, 50, 84, 112, 151</comment>).</mixed-citation></ref>
<ref id="CIT175"><label>[IS15]</label> <mixed-citation publication-type="book"><target target-type="page" id="pges_237"/><person-group person-group-type="author"><string-name><given-names>Sergey</given-names> <surname>Ioffe</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Christian</given-names> <surname>Szegedy</surname></string-name></person-group>. <chapter-title>&#x201C;Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2015</year>, <comment>pp.</comment> <fpage>448</fpage>&#x2013;<lpage>456</lpage> (<comment>cit. on pp. 54, 64, 194</comment>).</mixed-citation></ref>
<ref id="CIT176"><label>[Iso+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Phillip</given-names> <surname>Isola</surname></string-name>, <string-name><given-names>Jun-Yan</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Tinghui</given-names> <surname>Zhou</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Alexei A.</given-names> <surname>Efros</surname></string-name></person-group>. <article-title>&#x201C;Image-to-Image Translation with Conditional Adversarial Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>5967</fpage>&#x2013;<lpage>5976</lpage> (<comment>cit. on pp. 38, 39, 54, 56, 78, 79, 94, 107, 128, 132, 193</comment>).</mixed-citation></ref>
<ref id="CIT177"><label>[Jak+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wenzel</given-names> <surname>Jakob</surname></string-name>, <string-name><given-names>S&#x00E9;bastien</given-names> <surname>Speierer</surname></string-name>, <string-name><given-names>Nicolas</given-names> <surname>Roussel</surname></string-name>, <string-name><given-names>Merlin</given-names> <surname>Nimier-David</surname></string-name>, <string-name><given-names>Delio</given-names> <surname>Vicini</surname></string-name>, <string-name><given-names>Tizian</given-names> <surname>Zeltner</surname></string-name>, <string-name><given-names>Baptiste</given-names> <surname>Nicolet</surname></string-name>, <string-name><given-names>Miguel</given-names> <surname>Crespo</surname></string-name>, <string-name><given-names>Vincent</given-names> <surname>Leroy</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ziyi</given-names> <surname>Zhang</surname></string-name></person-group>. <source><italic>Mitsuba 3 Renderer</italic></source>. <comment>Version 3.1.1</comment>. <uri xlink:href="https://mitsuba-renderer.org">https://mitsuba-renderer.org</uri>. <year>2022</year> (<comment>cit. on pp. 189, 202</comment>).</mixed-citation></ref>
<ref id="CIT178"><label>[Jam+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ond&#x0159;ej</given-names> <surname>Jamri&#x0161;ka</surname></string-name>, <string-name><given-names>&#x0160;&#x00E1;rka</given-names> <surname>Sochorov&#x00E1;</surname></string-name>, <string-name><given-names>Ond&#x0159;ej</given-names> <surname>Texler</surname></string-name>, <string-name><given-names>Michal</given-names> <surname>Luk&#x00E1;&#x010D;</surname></string-name>, <string-name><given-names>Jakub</given-names> <surname>Fi&#x0161;er</surname></string-name>, <string-name><given-names>Jingwan</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Sy&#x0060;kora</surname></string-name></person-group>. <article-title>&#x201C;Stylizing Video by Example&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>38</volume>.<issue>4</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT179"><label>[Jan+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Michael</given-names> <surname>Janner</surname></string-name>, <string-name><given-names>Jiajun</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Tejas D</given-names> <surname>Kulkarni</surname></string-name>, <string-name><given-names>Ilker</given-names> <surname>Yildirim</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Josh</given-names> <surname>Tenenbaum</surname></string-name></person-group>. <article-title>&#x201C;Self-Supervised Intrinsic Image Decomposition&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>5936</fpage>&#x2013;<lpage>5946</lpage> (<comment>cit. on pp. 64, 67, 69, 78</comment>).</mixed-citation></ref>
<ref id="CIT180"><label>[JBH19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Miguel</given-names> <surname>Jaques</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Burke</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Timothy</given-names> <surname>Hospedales</surname></string-name></person-group>. <article-title>&#x201C;Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2019</year> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT181"><label>[JCJ09]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wojciech</given-names> <surname>Jarosz</surname></string-name>, <string-name><given-names>Nathan A.</given-names> <surname>Carr</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Henrik</given-names> <surname>Wann Jensen</surname></string-name></person-group>. <article-title>&#x201C;Importance Sampling Spherical Harmonics&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum (Proceedings of Eurographics)</italic></source> <volume>28</volume>.<issue>2</issue> (<year>2009</year>), <comment>pp.</comment> <fpage>577</fpage>&#x2013;<lpage>586</lpage> (<comment>cit. on p. 185</comment>).</mixed-citation></ref>
<ref id="CIT182"><label>[JBV16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Nikolay</given-names> <surname>Jetchev</surname></string-name>, <string-name><given-names>Urs</given-names> <surname>Bergmann</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Roland</given-names> <surname>Vollgraf</surname></string-name></person-group>. <article-title>&#x201C;Texture Synthesis with Spatial Generative Adversarial Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1611.08207</italic></source> (<year>2016</year>) (<comment>cit. on pp. 72, 74</comment>).</mixed-citation></ref>
<ref id="CIT183"><label>[Jia+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Liming</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>Bo</given-names> <surname>Dai</surname></string-name>, <string-name><given-names>Wayne</given-names> <surname>Wu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Chen</given-names> <surname>Change Loy</surname></string-name></person-group>. <article-title>&#x201C;Focal Frequency Loss for Image Reconstruction and Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>13919</fpage>&#x2013;<lpage>13929</lpage> (<comment>cit. on pp. 107, 128</comment>).</mixed-citation></ref>
<ref id="CIT184"><label>[Jin+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wenhua</given-names> <surname>Jin</surname></string-name>, <string-name><given-names>Beibei</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name>, <string-name><given-names>Yu</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Steve</given-names> <surname>Marschner</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ling-Qi</given-names> <surname>Yan</surname></string-name></person-group>. <article-title>&#x201C;Woven Fabric Capture from a <target target-type="page" id="pges_238"/>Single Photo&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of SIGGRAPH Asia</italic></source> <volume>2022</volume>. <year>2022</year> (<comment>cit. on p. 125</comment>).</mixed-citation></ref>
<ref id="CIT185"><label>[Jin+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yongcheng</given-names> <surname>Jing</surname></string-name>, <string-name><given-names>Yezhou</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Zunlei</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>Jingwen</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>Yizhou</given-names> <surname>Yu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mingli</given-names> <surname>Song</surname></string-name></person-group>. <article-title>&#x201C;Neural Style Transfer: A review&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Visualization and Computer Graphics</italic></source> (<year>2019</year>) (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT186"><label>[JAF16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Justin</given-names> <surname>Johnson</surname></string-name>, <string-name><given-names>Alexandre</given-names> <surname>Alahi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Li</given-names> <surname>Fei-Fei</surname></string-name></person-group>. <article-title>&#x201C;Perceptual Losses for Real-time Style Transfer and Super-resolution&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <year>2016</year>, <comment>pp.</comment> <fpage>694</fpage>&#x2013;<lpage>711</lpage> (<comment>cit. on pp. 32, 74</comment>).</mixed-citation></ref>
<ref id="CIT187"><label>[Jos+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Laurent</given-names> <surname>Valentin Jospin</surname></string-name>, <string-name><given-names>Hamid</given-names> <surname>Laga</surname></string-name>, <string-name><given-names>Farid</given-names> <surname>Boussaid</surname></string-name>, <string-name><given-names>Wray</given-names> <surname>Buntine</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mohammed</given-names> <surname>Bennamoun</surname></string-name></person-group>. <article-title>&#x201C;Hands-on Bayesian Neural Networks&#x2014;A Tutorial for Deep Learning Users&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Computational Intelligence Magazine</italic></source> <volume>17</volume>.<issue>2</issue> (<year>2022</year>), <comment>pp.</comment> <fpage>29</fpage>&#x2013;<lpage>48</lpage> (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT188"><label>[JC20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Eunjung</given-names> <surname>Ju</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Myung</given-names> <surname>Geol Choi</surname></string-name></person-group>. <article-title>&#x201C;Estimating Cloth Simulation Parameters From a Static Drape Using Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Access</italic></source> <volume>8</volume> (<year>2020</year>), <comment>pp.</comment> <fpage>195113</fpage>&#x2013;<lpage>195121</lpage> (<comment>cit. on pp. 156, 159</comment>).</mixed-citation></ref>
<ref id="CIT189"><label>[KBH03]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Florian</given-names> <surname>Kainz</surname></string-name>, <string-name><given-names>Rod</given-names> <surname>Bogart</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Drew</given-names> <surname>Hess</surname></string-name></person-group>. <article-title>&#x201C;The OpenEXR Image File Format&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM SIGGRAPH Technical Sketches</italic></source> (<year>2003</year>) (<comment>cit. on p. 185</comment>).</mixed-citation></ref>
<ref id="CIT190"><label>[Kaj86]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>James T</given-names> <surname>Kajiya</surname></string-name></person-group>. <article-title>&#x201C;The Rendering Equation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 13th annual conference on Computer graphics and interactive techniques</italic></source>. <year>1986</year>, <comment>pp.</comment> <fpage>143</fpage>&#x2013;<lpage>150</lpage> (<comment>cit. on pp. 4, 8</comment>).</mixed-citation></ref>
<ref id="CIT191"><label>[Kam+16]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Christos</given-names> <surname>Kampouris</surname></string-name>, <string-name><given-names>Stefanos</given-names> <surname>Zafeiriou</surname></string-name>, <string-name><given-names>Abhijeet</given-names> <surname>Ghosh</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Sotiris</given-names> <surname>Malassiotis</surname></string-name></person-group>. <article-title>&#x201C;Fine-Grained Material Classification Using Micro-Geometry and Reflectance&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2016</year>, <comment>pp.</comment> <fpage>778</fpage>&#x2013;<lpage>792</lpage> (<comment>cit. on p. 40</comment>).</mixed-citation></ref>
<ref id="CIT192"><label>[KG13]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Brian</given-names> <surname>Karis</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Epic</given-names> <surname>Games</surname></string-name></person-group>. <article-title>&#x201C;Real Shading in Unreal Engine 4&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proc. Physically Based Shading Theory Practice</italic></source> <volume>4</volume>.<issue>3</issue> (<year>2013</year>), <comment>p. 1</comment> (<comment>cit. on p. 127</comment>).</mixed-citation></ref>
<ref id="CIT193"><label>[Kar+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tero</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>Timo</given-names> <surname>Aila</surname></string-name>, <string-name><given-names>Samuli</given-names> <surname>Laine</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jaakko</given-names> <surname>Lehtinen</surname></string-name></person-group>. <article-title>&#x201C;Progressive Growing of GANs for Improved Quality, Stability, and Variation&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2018</year> (<comment>cit. on pp. 74, 80</comment>).</mixed-citation></ref>
<ref id="CIT194"><label>[Kar+20a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tero</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>Miika</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>Janne</given-names> <surname>Hellsten</surname></string-name>, <string-name><given-names>Samuli</given-names> <surname>Laine</surname></string-name>, <string-name><given-names>Jaakko</given-names> <surname>Lehtinen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Timo</given-names> <surname>Aila</surname></string-name></person-group>. <article-title>&#x201C;Training Generative Adversarial Networks with Limited Data&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2006.06676</italic></source> (<year>2020</year>) (<comment>cit. on pp. 36, 109</comment>).</mixed-citation></ref>
<ref id="CIT195"><label>[Kar+21]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_239"/><person-group person-group-type="author"><string-name><given-names>Tero</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>Miika</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>Samuli</given-names> <surname>Laine</surname></string-name>, <string-name><given-names>Erik</given-names> <surname>H&#x00E4;rk&#x00F6;nen</surname></string-name>, <string-name><given-names>Janne</given-names> <surname>Hellsten</surname></string-name>, <string-name><given-names>Jaakko</given-names> <surname>Lehtinen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Timo</given-names> <surname>Aila</surname></string-name></person-group>. <article-title>&#x201C;Alias-free Generative Adversarial Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <volume>Vol. 34</volume>. <year>2021</year>, <comment>pp.</comment> <fpage>852</fpage>&#x2013;<lpage>863</lpage> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT196"><label>[KLA19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tero</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>Samuli</given-names> <surname>Laine</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Timo</given-names> <surname>Aila</surname></string-name></person-group>. <article-title>&#x201C;A Style-Based Generator Architecture for Generative Adversarial Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>4401</fpage>&#x2013;<lpage>4410</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT197"><label>[Kar+20b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tero</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>Samuli</given-names> <surname>Laine</surname></string-name>, <string-name><given-names>Miika</given-names> <surname>Aittala</surname></string-name>, <string-name><given-names>Janne</given-names> <surname>Hellsten</surname></string-name>, <string-name><given-names>Jaakko</given-names> <surname>Lehtinen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Timo</given-names> <surname>Aila</surname></string-name></person-group>. <article-title>&#x201C;Analyzing and Improving the Image Quality of StyleGAN&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <month>June</month> <year>2020</year> (<comment>cit. on pp. 74, 85</comment>).</mixed-citation></ref>
<ref id="CIT198"><label>[Kas+15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexandre</given-names> <surname>Kaspar</surname></string-name>, <string-name><given-names>Boris</given-names> <surname>Neubert</surname></string-name>, <string-name><given-names>Dani</given-names> <surname>Lischinski</surname></string-name>, <string-name><given-names>Mark</given-names> <surname>Pauly</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Johannes</given-names> <surname>Kopf</surname></string-name></person-group>. <article-title>&#x201C;Self Tuning Texture Optimization&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 34</volume>. <issue>2</issue>. <year>2015</year>, <comment>pp.</comment> <fpage>349</fpage>&#x2013;<lpage>359</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT199"><label>[KZP19]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Sergey</given-names> <surname>Kastryulin</surname></string-name>, <string-name><given-names>Dzhamil</given-names> <surname>Zakirov</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Denis</given-names> <surname>Prokopenko</surname></string-name></person-group>. <source><italic>PyTorch Image Quality: Metrics and Measure for Image Quality Assessment</italic></source>. <publisher-name>Open-source software</publisher-name> <comment>available at</comment> <uri xlink:href="https://github.com/photosynthesis-team/piq">https://github.com/photosynthesis-team/piq</uri>. <year>2019</year> (<comment>cit. on pp. 182, 198</comment>).</mixed-citation></ref>
<ref id="CIT200"><label>[Kau17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Eric</given-names> <surname>Kauderer-Abrams</surname></string-name></person-group>. <article-title>&#x201C;Quantifying Translation-invariance in Convolutional Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1801.01450</italic></source> (<year>2017</year>) (<comment>cit. on p. 37</comment>).</mixed-citation></ref>
<ref id="CIT201"><label>[Kau+00]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Kautz</surname></string-name>, <string-name><given-names>Pere-Pau</given-names> <surname>V&#x00E1;zquez</surname></string-name>, <string-name><given-names>Wolfgang</given-names> <surname>Heidrich</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hans-Peter</given-names> <surname>Seidel</surname></string-name></person-group>. <article-title>&#x201C;Unified Approach to Prefiltered Environment Maps&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the Eurographics Workshop on Rendering Techniques</italic></source> <volume>2000</volume>. <year>2000</year>, <comment>pp.</comment> <fpage>185</fpage>&#x2013;<lpage>196</lpage> (<comment>cit. on p. 185</comment>).</mixed-citation></ref>
<ref id="CIT202"><label>[Kaw80]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sueo</given-names> <surname>Kawabata</surname></string-name></person-group>. <article-title>&#x201C;The Standardization and Analysis of Hand Evaluation&#x201D;</article-title>. <comment>In</comment>: <source><italic>The Textile Machinery Society of Japan</italic></source> (<year>1980</year>) (<comment>cit. on pp. 156, 158</comment>).</mixed-citation></ref>
<ref id="CIT203"><label>[KG17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alex</given-names> <surname>Kendall</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yarin</given-names> <surname>Gal</surname></string-name></person-group>. <article-title>&#x201C;What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?&#x201D;</article-title> <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>30</volume> (<year>2017</year>) (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT204"><label>[Khe+21]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Ilyes</given-names> <surname>Khemakhem</surname></string-name>, <string-name><given-names>Ricardo</given-names> <surname>Monti</surname></string-name>, <string-name><given-names>Robert</given-names> <surname>Leech</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Aapo</given-names> <surname>Hyvarinen</surname></string-name></person-group>. <article-title>&#x201C;Causal Autoregressive Flows&#x201D;</article-title>. <comment>In</comment>: <source><italic>International <target target-type="page" id="pges_240"/>Conference on Artificial Intelligence and Statistics</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>3520</fpage>&#x2013;<lpage>3528</lpage> (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT205"><label>[KB15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Diederik P.</given-names> <surname>Kingma</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jimmy</given-names> <surname>Ba</surname></string-name></person-group>. <article-title>&#x201C;Adam: A Method for Stochastic Optimization&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2015</year> (<comment>cit. on pp. 39, 55, 67, 83, 94, 110, 142, 194</comment>).</mixed-citation></ref>
<ref id="CIT206"><label>[Kin+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Durk P</given-names> <surname>Kingma</surname></string-name>, <string-name><given-names>Tim</given-names> <surname>Salimans</surname></string-name>, <string-name><given-names>Rafal</given-names> <surname>Jozefowicz</surname></string-name>, <string-name><given-names>Xi</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Ilya</given-names> <surname>Sutskever</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Max</given-names> <surname>Welling</surname></string-name></person-group>. <article-title>&#x201C;Improved Variational Inference with Inverse Autoregressive Flow&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>29</volume> (<year>2016</year>) (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT207"><label>[KD18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Durk P.</given-names> <surname>Kingma</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Prafulla</given-names> <surname>Dhariwal</surname></string-name></person-group>. <article-title>&#x201C;Glow: Generative Flow with Invertible 1x1 Convolutions&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <volume>Vol. 31.</volume> <year>2018</year> (<comment>cit. on pp. 186, 191</comment>).</mixed-citation></ref>
<ref id="CIT208"><label>[Kir+23]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexander</given-names> <surname>Kirillov</surname></string-name>, <string-name><given-names>Eric</given-names> <surname>Mintun</surname></string-name>, <string-name><given-names>Nikhila</given-names> <surname>Ravi</surname></string-name>, <string-name><given-names>Hanzi</given-names> <surname>Mao</surname></string-name>, <string-name><given-names>Chloe</given-names> <surname>Rolland</surname></string-name>, <string-name><given-names>Laura</given-names> <surname>Gustafson</surname></string-name>, <string-name><given-names>Tete</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>Spencer</given-names> <surname>Whitehead</surname></string-name>, <string-name><given-names>Alexander C</given-names> <surname>Berg</surname></string-name>, <string-name><given-names>Wan-Yen</given-names> <surname>Lo</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;Segment Anything&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2304.02643</italic></source> (<year>2023</year>) (<comment>cit. on p. 212</comment>).</mixed-citation></ref>
<ref id="CIT209"><label>[KPB20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ivan</given-names> <surname>Kobyzev</surname></string-name>, <string-name><given-names>Simon JD</given-names> <surname>Prince</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Marcus A</given-names> <surname>Brubaker</surname></string-name></person-group>. <article-title>&#x201C;Normalizing Flows: An Introduction and Review of Current Methods&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Pattern Analysis and Machine Intelligence</italic></source> <volume>43</volume>.<issue>11</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>3964</fpage>&#x2013;<lpage>3979</lpage> (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT210"><label>[Kol+20]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Alexander</given-names> <surname>Kolesnikov</surname></string-name>, <string-name><given-names>Lucas</given-names> <surname>Beyer</surname></string-name>, <string-name><given-names>Xiaohua</given-names> <surname>Zhai</surname></string-name>, <string-name><given-names>Joan</given-names> <surname>Puigcerver</surname></string-name>, <string-name><given-names>Jessica</given-names> <surname>Yung</surname></string-name>, <string-name><given-names>Sylvain</given-names> <surname>Gelly</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Neil</given-names> <surname>Houlsby</surname></string-name></person-group>. <chapter-title>&#x201C;Big Transfer (bit): General Visual Representation Learning&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>491</fpage>&#x2013;<lpage>507</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT211"><label>[KSL19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Simon</given-names> <surname>Kornblith</surname></string-name>, <string-name><given-names>Jonathon</given-names> <surname>Shlens</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Quoc V</given-names> <surname>Le</surname></string-name></person-group>. <article-title>&#x201C;Do Better Imagenet Models Transfer Better?&#x201D;</article-title> <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>2661</fpage>&#x2013;<lpage>2671</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT212"><label>[Kou+03]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Melissa L</given-names> <surname>Koudelka</surname></string-name>, <string-name><given-names>Sebastian</given-names> <surname>Magda</surname></string-name>, <string-name><given-names>Peter N</given-names> <surname>Belhumeur</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>David J</given-names> <surname>Kriegman</surname></string-name></person-group>. <article-title>&#x201C;Acquisition, Compression, and Synthesis of Bidirectional Texture Functions&#x201D;</article-title>. <comment>In</comment>: <source><italic>3rd International Workshop on Texture Analysis and Synthesis (Texture 2003)</italic></source>. <year>2003</year>, <comment>pp.</comment> <fpage>59</fpage>&#x2013;<lpage>64</lpage> (<comment>cit. on p. 103</comment>).</mixed-citation></ref>
<ref id="CIT213"><label>[Kry+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Michael C</given-names> <surname>Krygier</surname></string-name>, <string-name><given-names>Tyler</given-names> <surname>LaBonte</surname></string-name>, <string-name><given-names>Carianne</given-names> <surname>Martinez</surname></string-name>, <string-name><given-names>Chance</given-names> <surname>Norris</surname></string-name>, <string-name><given-names>Krish</given-names> <surname>Sharma</surname></string-name>, <string-name><given-names>Lincoln N</given-names> <surname>Collins</surname></string-name>, <string-name><given-names>Partha P</given-names> <surname>Mukherjee</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Scott A</given-names> <surname>Roberts</surname></string-name></person-group>. <article-title>&#x201C;Quantifying the Unknown Impact of Segmentation Uncertainty on Image-Based <target target-type="page" id="pges_241"/>Simulations&#x201D;</article-title>. <comment>In</comment>: <source><italic>Nature Communications</italic></source> <volume>12</volume>.<issue>1</issue> (<year>2021</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT214"><label>[KLG20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sandra</given-names> <surname>Kuijpers</surname></string-name>, <string-name><given-names>Christiane</given-names> <surname>Luible-B&#x00E4;r</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Hugh Gong</surname></string-name></person-group>. <article-title>&#x201C;The Measurement of Fabric Properties for Virtual Simulation&#x2014;A Critical Review&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE SA INDUSTRY CONNECTIONS</italic></source> (<year>2020</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>43</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT215"><label>[KL51]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Solomon</given-names> <surname>Kullback</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Richard A</given-names> <surname>Leibler</surname></string-name></person-group>. <article-title>&#x201C;On Information and Sufficiency&#x201D;</article-title>. <comment>In</comment>: <source><italic>The Annals of Mathematical Statistics</italic></source> <volume>22</volume>.<issue>1</issue> (<year>1951</year>), <comment>pp.</comment> <fpage>79</fpage>&#x2013;<lpage>86</lpage> (<comment>cit. on pp. 196, 197</comment>).</mixed-citation></ref>
<ref id="CIT216"><label>[Kum+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ananya</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>Aditi</given-names> <surname>Raghunathan</surname></string-name>, <string-name><given-names>Robbie</given-names> <surname>Jones</surname></string-name>, <string-name><given-names>Tengyu</given-names> <surname>Ma</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Percy</given-names> <surname>Liang</surname></string-name></person-group>. <article-title>&#x201C;Fine-tuning Can Distort Pretrained Features and Underperform Out-of-distribution&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2202.10054</italic></source> (<year>2022</year>) (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT217"><label>[Kum+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Manoj</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>Mohammad</given-names> <surname>Babaeizadeh</surname></string-name>, <string-name><given-names>Dumitru</given-names> <surname>Erhan</surname></string-name>, <string-name><given-names>Chelsea</given-names> <surname>Finn</surname></string-name>, <string-name><given-names>Sergey</given-names> <surname>Levine</surname></string-name>, <string-name><given-names>Laurent</given-names> <surname>Dinh</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Durk</given-names> <surname>Kingma</surname></string-name></person-group>. <article-title>&#x201C;Videoflow: A Flow-based Generative Model for Video&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1903.01434</italic></source> <volume>2</volume>.<issue>5</issue> (<year>2019</year>), <comment>p.</comment> <fpage>3</fpage> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT218"><label>[Kur+19]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Karol</given-names> <surname>Kurach</surname></string-name>, <string-name><given-names>Mario</given-names> <surname>Lu&#x010D;i&#x0107;</surname></string-name>, <string-name><given-names>Xiaohua</given-names> <surname>Zhai</surname></string-name>, <string-name><given-names>Marcin</given-names> <surname>Michalski</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Sylvain</given-names> <surname>Gelly</surname></string-name></person-group>. <chapter-title>&#x201C;A Large-Scale Study on Regularization and Normalization in GANs&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>3581</fpage>&#x2013;<lpage>3590</lpage> (<comment>cit. on pp. 77, 94</comment>).</mixed-citation></ref>
<ref id="CIT219"><label>[Kur+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexander</given-names> <surname>Kurz</surname></string-name>, <string-name><given-names>Katja</given-names> <surname>Hauser</surname></string-name>, <string-name><given-names>Hendrik</given-names> <surname>Alexander Mehrtens</surname></string-name>, <string-name><given-names>Eva</given-names> <surname>Krieghoff-Henning</surname></string-name>, <string-name><given-names>Achim</given-names> <surname>Hekler</surname></string-name>, <string-name><given-names>Jakob</given-names> <surname>Nikolas Kather</surname></string-name>, <string-name><given-names>Stefan</given-names> <surname>Fr&#x00F6;hling</surname></string-name>, <string-name><given-names>Christof von</given-names> <surname>Kalle</surname></string-name>, <string-name><given-names>Titus</given-names> <surname>Josef Brinker</surname></string-name>,</person-group> <etal>et al.</etal> <article-title>&#x201C;Uncertainty Estimation in Medical Image Classification: Systematic Review&#x201D;</article-title>. <comment>In</comment>: <source><italic>JMIR Medical Informatics</italic></source> <volume>10</volume>.<issue>8</issue> (<year>2022</year>) (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT220"><label>[Kuz+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexandr</given-names> <surname>Kuznetsov</surname></string-name>, <string-name><given-names>Krishna</given-names> <surname>Mullia</surname></string-name>, <string-name><given-names>Zexiang</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name></person-group>. <article-title>&#x201C;NeuMIP: Multi-Resolution Neural Materials&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>40.</volume><issue>4</issue> (<year>2021</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>13</lpage> (<comment>cit. on pp. 18, 102, 103, 106, 107, 109&#x2013;116, 119, 120, 187, 207, 259, 264</comment>).</mixed-citation></ref>
<ref id="CIT221"><label>[Kuz+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexandr</given-names> <surname>Kuznetsov</surname></string-name>, <string-name><given-names>Xuezheng</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Krishna</given-names> <surname>Mullia</surname></string-name>, <string-name><given-names>Fujun</given-names> <surname>Luan</surname></string-name>, <string-name><given-names>Zexiang</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Milos</given-names> <surname>Hasan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name></person-group>. <article-title>&#x201C;Rendering Neural Materials on Curved Surfaces&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM SIGGRAPH 2022 Conference Proceedings</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>9</lpage> (<comment>cit. on pp. 102, 104, 116, 187</comment>).</mixed-citation></ref>
<ref id="CIT222"><label>[Kwa+05]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Vivek</given-names> <surname>Kwatra</surname></string-name>, <string-name><given-names>Irfan</given-names> <surname>Essa</surname></string-name>, <string-name><given-names>Aaron</given-names> <surname>Bobick</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Nipun</given-names> <surname>Kwatra</surname></string-name></person-group>. <chapter-title>&#x201C;Texture Optimization for Example-Based Synthesis&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source>. <volume>Vol. 24. 3</volume>. <publisher-name>ACM</publisher-name>. <year>2005</year>, <comment>pp.</comment> <fpage>795</fpage>&#x2013;<lpage>802</lpage> (<comment>cit. on pp. 72, 74</comment>).</mixed-citation></ref>
<ref id="CIT223"><label>[Kwa+03]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_242"/><person-group person-group-type="author"><string-name><given-names>Vivek</given-names> <surname>Kwatra</surname></string-name>, <string-name><given-names>Arno</given-names> <surname>Sch&#x00F6;dl</surname></string-name>, <string-name><given-names>Irfan</given-names> <surname>Essa</surname></string-name>, <string-name><given-names>Greg</given-names> <surname>Turk</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Aaron</given-names> <surname>Bobick</surname></string-name></person-group>. <article-title>&#x201C;Graphcut Textures: Image and Video Synthesis Ysing Graph Cuts&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>22</volume>.<issue>3</issue> (<year>2003</year>), <comment>pp.</comment> <fpage>277</fpage>&#x2013;<lpage>286</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT224"><label>[LGG19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Manuel</given-names> <surname>Lagunas</surname></string-name>, <string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Diego</given-names> <surname>Gutierrez</surname></string-name></person-group>. <article-title>&#x201C;Learning Icons Appearance Similarity&#x201D;</article-title>. <comment>In</comment>: <source><italic>Multimedia Tools and Applications</italic></source> <volume>78</volume>.<issue>8</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>10733</fpage>&#x2013;<lpage>10751</lpage> (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT225"><label>[Lag+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Manuel</given-names> <surname>Lagunas</surname></string-name>, <string-name><given-names>Sandra</given-names> <surname>Malpica</surname></string-name>, <string-name><given-names>Ana</given-names> <surname>Serrano</surname></string-name>, <string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name>, <string-name><given-names>Diego</given-names> <surname>Gutierrez</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Belen</given-names> <surname>Masia</surname></string-name></person-group>. <article-title>&#x201C;A Similarity Measure for Material Appearance&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>38</volume>.<issue>4</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>12</lpage> (<comment>cit. on pp. 157, 160, 180</comment>).</mixed-citation></ref>
<ref id="CIT226"><label>[LKA13]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Samuli</given-names> <surname>Laine</surname></string-name>, <string-name><given-names>Tero</given-names> <surname>Karras</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Timo</given-names> <surname>Aila</surname></string-name></person-group>. <article-title>&#x201C;Megakernels Considered Harmful: Wavefront Path Tracing on GPUs&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 5th High-Performance Graphics Conference</italic></source>. <year>2013</year>, <comment>pp.</comment> <fpage>137</fpage>&#x2013;<lpage>143</lpage> (<comment>cit. on pp. 109, 195</comment>).</mixed-citation></ref>
<ref id="CIT227"><label>[LPB17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Balaji</given-names> <surname>Lakshminarayanan</surname></string-name>, <string-name><given-names>Alexander</given-names> <surname>Pritzel</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Charles</given-names> <surname>Blundell</surname></string-name></person-group>. <article-title>&#x201C;Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>30</volume> (<year>2017</year>) (<comment>cit. on pp. 126, 130</comment>).</mixed-citation></ref>
<ref id="CIT228"><label>[LS98]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Greg</given-names> <surname>Ward Larson</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Rob</given-names> <surname>Shakespeare</surname></string-name></person-group>. <source><italic>Rendering with Radiance: The Art and Science of Lighting Visualization</italic></source>. <publisher-loc>San Francisco, CA, USA</publisher-loc>: <publisher-name>Morgan Kaufmann Publishers Inc.</publisher-name>, <year>1998</year> (<comment>cit. on p. 185</comment>).</mixed-citation></ref>
<ref id="CIT229"><label>[Lav+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Guillaume</given-names> <surname>Lavou&#x00E9;</surname></string-name>, <string-name><given-names>Nicolas</given-names> <surname>Bonneel</surname></string-name>, <string-name><given-names>Jean-Philippe</given-names> <surname>Farrugia</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Cyril</given-names> <surname>Soler</surname></string-name></person-group>. <article-title>&#x201C;Perceptual Quality of BRDF Approximations: Dataset and Metrics&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 40</volume>. <issue>2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>327</fpage>&#x2013;<lpage>338</lpage> (<comment>cit. on pp. 130, 131, 193</comment>).</mixed-citation></ref>
<ref id="CIT230"><label>[Led+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Christian</given-names> <surname>Ledig</surname></string-name>, <string-name><given-names>Lucas</given-names> <surname>Theis</surname></string-name>, <string-name><given-names>Ferenc</given-names> <surname>Husz&#x00E1;r</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source> <comment>2017-Janua</comment> (<month>Sept</month>. <year>2017</year>), <comment>pp.</comment> <fpage>105</fpage>&#x2013;<lpage>114</lpage> (<comment>cit. on pp. 79, 94</comment>).</mixed-citation></ref>
<ref id="CIT231"><label>[LH06]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sylvain</given-names> <surname>Lefebvre</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hugues</given-names> <surname>Hoppe</surname></string-name></person-group>. <article-title>&#x201C;Appearance-Space Texture Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>25</volume>.<issue>3</issue> (<year>2006</year>), <comment>pp.</comment> <fpage>541</fpage>&#x2013;<lpage>548</lpage> (<comment>cit. on pp. 34, 104</comment>).</mixed-citation></ref>
<ref id="CIT232"><label>[Lei+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhao</given-names> <surname>Lei</surname></string-name>, <string-name><given-names>Yi</given-names> <surname>Zeng</surname></string-name>, <string-name><given-names>Peng</given-names> <surname>Liu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Xiaohui</given-names> <surname>Su</surname></string-name></person-group>. <article-title>&#x201C;Active Deep Learning for Hyperspectral Image Classification with Uncertainty Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Geoscience and Remote Sensing Letters</italic></source> <volume>19</volume> (<year>2021</year>) (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT233"><label>[LM01]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_243"/><person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>Leung</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jitendra</given-names> <surname>Malik</surname></string-name></person-group>. <article-title>&#x201C;Representing and Recognizing the Visual Appearance of Materials Using Three-Dimensional Textons&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Journal of Computer Vision</italic></source> <volume>43</volume>.<issue>1</issue> (<year>2001</year>), <comment>pp.</comment> <fpage>29</fpage>&#x2013;<lpage>44</lpage> (<comment>cit. on pp. 33, 39</comment>).</mixed-citation></ref>
<ref id="CIT234"><label>[LLB22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wei-Hong</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Xialei</given-names> <surname>Liu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hakan</given-names> <surname>Bilen</surname></string-name></person-group>. <article-title>&#x201C;Cross-domain Few-shot Learning with Task-specific Adapters&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>7161</fpage>&#x2013;<lpage>7170</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT235"><label>[Li+17a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xiao</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Yue</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Pieter</given-names> <surname>Peers</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Xin</given-names> <surname>Tong</surname></string-name></person-group>. <article-title>&#x201C;Modeling Surface Appearance from a Single Photograph Using Self-Augmented Convolutional Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>36</volume>.<issue>4</issue> (<year>2017</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on pp. 34, 124</comment>).</mixed-citation></ref>
<ref id="CIT236"><label>[Li+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xiaoyu</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Bo</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Pedro V</given-names> <surname>Sander</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name></person-group>. <article-title>&#x201C;Blind Geometric Distortion Correction on Images through Deep Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>4855</fpage>&#x2013;<lpage>4864</lpage> (<comment>cit. on p. 93</comment>).</mixed-citation></ref>
<ref id="CIT237"><label>[Li+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yifei</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Tao</given-names> <surname>Du</surname></string-name>, <string-name><given-names>Kui</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Jie</given-names> <surname>Xu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Wojciech</given-names> <surname>Matusik</surname></string-name></person-group>. <article-title>&#x201C;DiffCloth: Differentiable Cloth Simulation with Dry Frictional Contact&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> (<year>2022</year>) (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT238"><label>[Li+17b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yijun</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Chen</given-names> <surname>Fang</surname></string-name>, <string-name><given-names>Jimei</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Zhaowen</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Xin</given-names> <surname>Lu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ming-Hsuan</given-names> <surname>Yang</surname></string-name></person-group>. <article-title>&#x201C;Universal Style Transfer via Feature Transforms&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>386</fpage>&#x2013;<lpage>396</lpage> (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT239"><label>[LAA08]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yuanzhen</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Edward</given-names> <surname>Adelson</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Aseem</given-names> <surname>Agarwala</surname></string-name></person-group>. <article-title>&#x201C;ScribbleBoost: Adding Classification to Edge-Aware Interpolation of Local Image and Video Adjustments&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 27</volume>. <issue>4</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2008</year>, <comment>pp.</comment> <fpage>1255</fpage>&#x2013;<lpage>1264</lpage> (<comment>cit. on p. 33</comment>).</mixed-citation></ref>
<ref id="CIT240"><label>[LS18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhengqi</given-names> <surname>Li</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Noah</given-names> <surname>Snavely</surname></string-name></person-group>. <article-title>&#x201C;Learning Intrinsic Image Decomposition From Watching the World&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <month>June</month> <year>2018</year> (<comment>cit. on p. 78</comment>).</mixed-citation></ref>
<ref id="CIT241"><label>[Li+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhengqin</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Mohammad</given-names> <surname>Shafiei</surname></string-name>, <string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name>, <string-name><given-names>Kalyan</given-names> <surname>Sunkavalli</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Man-mohan</given-names> <surname>Chandraker</surname></string-name></person-group>. <article-title>&#x201C;Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF From a Single Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>2475</fpage>&#x2013;<lpage>2484</lpage> (<comment>cit. on pp. 72, 74, 87&#x2013;91, 95, 99, 114, 203</comment>).</mixed-citation></ref>
<ref id="CIT242"><label>[LSC18]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_244"/><person-group person-group-type="author"><string-name><given-names>Zhengqin</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Kalyan</given-names> <surname>Sunkavalli</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Manmohan</given-names> <surname>Chandraker</surname></string-name></person-group>. <article-title>&#x201C;Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <year>2018</year>, <comment>pp.</comment> <fpage>72</fpage>&#x2013;<lpage>87</lpage> (<comment>cit. on pp. 72, 125</comment>).</mixed-citation></ref>
<ref id="CIT243"><label>[Li+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhengqin</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Zexiang</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name>, <string-name><given-names>Kalyan</given-names> <surname>Sunkavalli</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Manmohan</given-names> <surname>Chandraker</surname></string-name></person-group>. <article-title>&#x201C;Learning to Reconstruct Shape and Spatially-varying Reflectance from a Single Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>37</volume>.<issue>6</issue> (<year>2018</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on pp. 125, 129</comment>).</mixed-citation></ref>
<ref id="CIT244"><label>[LLK19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Junbang</given-names> <surname>Liang</surname></string-name>, <string-name><given-names>Ming</given-names> <surname>Lin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Vladlen</given-names> <surname>Koltun</surname></string-name></person-group>. <article-title>&#x201C;Differentiable Cloth Simulation for Inverse Problems&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>32</volume> (<year>2019</year>) (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT245"><label>[Lia+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>Yuan</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>Lu</given-names> <surname>Yuan</surname></string-name>, <string-name><given-names>Gang</given-names> <surname>Hua</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Sing</given-names> <surname>Bing Kang</surname></string-name></person-group>. <article-title>&#x201C;Visual Attribute Transfer through Deep Image Analogy&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>36</volume>.<issue>4</issue> (<year>2017</year>), <fpage>pp.</fpage> <fpage>1</fpage>&#x2013;<lpage>15</lpage> (<comment>cit. on pp. 31, 32, 46, 47, 56, 57, 60</comment>).</mixed-citation></ref>
<ref id="CIT246"><label>[Lic+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Lichy</surname></string-name>, <string-name><given-names>Jiaye</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Soumyadip</given-names> <surname>Sengupta</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>David W</given-names> <surname>Jacobs</surname></string-name></person-group>. <article-title>&#x201C;Shape and Material Capture at Home&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>6123</fpage>&#x2013;<lpage>6133</lpage> (<comment>cit. on p. 129</comment>).</mixed-citation></ref>
<ref id="CIT247"><label>[LPG19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yiming</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Pieter</given-names> <surname>Peers</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Abhijeet</given-names> <surname>Ghosh</surname></string-name></person-group>. <article-title>&#x201C;On-Site Example-Based Material Appearance Acquisition&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 38</volume>. <issue>4</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>15</fpage>&#x2013;<lpage>25</lpage> (<comment>cit. on pp. 34, 105</comment>).</mixed-citation></ref>
<ref id="CIT248"><label>[Liu+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Guilin</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Rohan</given-names> <surname>Taori</surname></string-name>, <string-name><given-names>Ting-Chun</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Zhiding</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>Shiqiu</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Fitsum A</given-names> <surname>Reda</surname></string-name>, <string-name><given-names>Karan</given-names> <surname>Sapra</surname></string-name>, <string-name><given-names>Andrew</given-names> <surname>Tao</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Bryan</given-names> <surname>Catanzaro</surname></string-name></person-group>. <article-title>&#x201C;Transposer: Universal Texture Synthesis Using Feature Maps as Transposed Convolution Filter&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2007.07243</italic></source> (<year>2020</year>) (<comment>cit. on pp. 72, 74, 75, 92</comment>).</mixed-citation></ref>
<ref id="CIT249"><label>[Liu+19a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Liyuan</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Haoming</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>Pengcheng</given-names> <surname>He</surname></string-name>, <string-name><given-names>Weizhu</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Xiaodong</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Jianfeng</given-names> <surname>Gao</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jiawei</given-names> <surname>Han</surname></string-name></person-group>. <article-title>&#x201C;On the Variance of the Adaptive Learning Rate and Beyond&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1908.03265</italic></source> (<year>2019</year>) (<comment>cit. on p. 110</comment>).</mixed-citation></ref>
<ref id="CIT250"><label>[Liu+19b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ming-Yu</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Xun</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Arun</given-names> <surname>Mallya</surname></string-name>, <string-name><given-names>Tero</given-names> <surname>Karras</surname></string-name>, <string-name><given-names>Timo</given-names> <surname>Aila</surname></string-name>, <string-name><given-names>Jaakko</given-names> <surname>Lehtinen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Kautz</surname></string-name></person-group>. <article-title>&#x201C;Few-shot Unsupervised Image-to-Image Translation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>10551</fpage>&#x2013;<lpage>10560</lpage> (<comment>cit. on p. 37</comment>).</mixed-citation></ref>
<ref id="CIT251"><label>[Liu+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Steven</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Xiuming</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Zhoutong</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Richard</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Jun-Yan</given-names> <surname>Zhu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Bryan</given-names> <surname>Russell</surname></string-name></person-group>. <article-title>&#x201C;Editing conditional <target target-type="page" id="pges_245"/>radiance fields&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>5773</fpage>&#x2013;<lpage>5783</lpage> (<comment>cit. on p. 116</comment>).</mixed-citation></ref>
<ref id="CIT252"><label>[Liu+22a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yuan</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Sida</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>Lingjie</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Qianqian</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Peng</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Christian</given-names> <surname>Theobalt</surname></string-name>, <string-name><given-names>Xiaowei</given-names> <surname>Zhou</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Wenping</given-names> <surname>Wang</surname></string-name></person-group>. <article-title>&#x201C;Neural Rays for Occlusion-aware Image-based Rendering&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>7824</fpage>&#x2013;<lpage>7833</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT253"><label>[Liu+22b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhuang</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Hanzi</given-names> <surname>Mao</surname></string-name>, <string-name><given-names>Chao-Yuan</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Christoph</given-names> <surname>Feichtenhofer</surname></string-name>, <string-name><given-names>Trevor</given-names> <surname>Darrell</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Saining</given-names> <surname>Xie</surname></string-name></person-group>. <article-title>&#x201C;A Convnet for the 2020s&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>11976</fpage>&#x2013;<lpage>11986</lpage> (<comment>cit. on pp. 108, 109, 119</comment>).</mixed-citation></ref>
<ref id="CIT254"><label>[LSD15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jonathan</given-names> <surname>Long</surname></string-name>, <string-name><given-names>Evan</given-names> <surname>Shelhamer</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Trevor</given-names> <surname>Darrell</surname></string-name></person-group>. <article-title>&#x201C;Fully Convolutional Networks for Semantic Segmentation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2015</year>, <comment>pp.</comment> <fpage>3431</fpage>&#x2013;<lpage>3440</lpage> (<comment>cit. on pp. 83, 94</comment>).</mixed-citation></ref>
<ref id="CIT255"><label>[LH17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ilya</given-names> <surname>Loshchilov</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Frank</given-names> <surname>Hutter</surname></string-name></person-group>. <article-title>&#x201C;Decoupled Weight Decay Regularization&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1711.05101</italic></source> (<year>2017</year>) (<comment>cit. on p. 176</comment>).</mixed-citation></ref>
<ref id="CIT256"><label>[Lua+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Fujun</given-names> <surname>Luan</surname></string-name>, <string-name><given-names>Shuang</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Kavita</given-names> <surname>Bala</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Zhao</given-names> <surname>Dong</surname></string-name></person-group>. <article-title>&#x201C;Unified Shape and SVBRDF Recovery Using Differentiable Monte Carlo Rendering&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 40</volume>. <issue>4</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>101</fpage>&#x2013;<lpage>113</lpage> (<comment>cit. on pp. 125, 129</comment>).</mixed-citation></ref>
<ref id="CIT257"><label>[Lug+20]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Andreas</given-names> <surname>Lugmayr</surname></string-name>, <string-name><given-names>Martin</given-names> <surname>Danelljan</surname></string-name>, <string-name><given-names>Luc</given-names> <surname>Van Gool</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Radu</given-names> <surname>Timofte</surname></string-name></person-group>. <chapter-title>&#x201C;Srflow: Learning the Super-resolution Space with Normalizing Flow&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>715</fpage>&#x2013;<lpage>732</lpage> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT258"><label>[LM08]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Christiane</given-names> <surname>Luible</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Nadia</given-names> <surname>Magnenat-Thalmann</surname></string-name></person-group>. <article-title>&#x201C;The Simulation of Cloth Using Accurate Physical Parameters&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the Tenth IASTED International Conference on Computer Graphics and Imaging (CGIM &#x2019;08)</italic></source>. <year>2008</year> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT259"><label>[LKS15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhaoliang</given-names> <surname>Lun</surname></string-name>, <string-name><given-names>Evangelos</given-names> <surname>Kalogerakis</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Alla</given-names> <surname>Sheffer</surname></string-name></person-group>. <article-title>&#x201C;Elements of Style: Learning Perceptual Shape Style Similarity&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on graphics (TOG)</italic></source> <volume>34</volume>.<issue>4</issue> (<year>2015</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>14</lpage> (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT260"><label>[MHN13]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Andrew L</given-names> <surname>Maas</surname></string-name>, <string-name><given-names>Awni Y</given-names> <surname>Hannun</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Andrew Y</given-names> <surname>Ng</surname></string-name></person-group>. <chapter-title>&#x201C;Rectifier Nonlinearities Improve Neural Network Acoustic Models&#x201D;</chapter-title>. <comment>In</comment>: <target target-type="page" id="pges_246"/><source><italic>International Conference on Machine Learning (ICML)</italic></source>. <volume>Vol. 30</volume>. <issue>1</issue>. <publisher-name>Citeseer</publisher-name>. <publisher-name>2013</publisher-name>, <fpage>p.</fpage> <fpage>3</fpage> (<comment>cit. on pp. 77, 94, 176</comment>).</mixed-citation></ref>
<ref id="CIT261"><label>[Mag+07]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Nadia</given-names> <surname>Magnenat-Thalmann</surname></string-name>, <string-name><given-names>Christiane</given-names> <surname>Luible</surname></string-name>, <string-name><given-names>Pascal</given-names> <surname>Volino</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Etienne</given-names> <surname>Lyard</surname></string-name></person-group>. <chapter-title>&#x201C;From Measured Fabric to the Simulation of Cloth&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>2007 10th IEEE International Conference on Computer-Aided Design and Computer Graphics</italic></source>. <publisher-name>IEEE</publisher-name>. <year>2007</year>, <comment>pp.</comment> <fpage>7</fpage>&#x2013;<lpage>18</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT262"><label>[MR10]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S&#x00E9;bastien</given-names> <surname>Marcel</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yann</given-names> <surname>Rodriguez</surname></string-name></person-group>. <article-title>&#x201C;Torchvision the Machine-vision Package of Torch&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 18th ACM International Conference on Multimedia</italic></source>. <year>2010</year>, <comment>pp.</comment> <fpage>1485</fpage>&#x2013;<lpage>1488</lpage> (<comment>cit. on pp. 67, 84, 110, 142, 194</comment>).</mixed-citation></ref>
<ref id="CIT263"><label>[Mar+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Morteza</given-names> <surname>Mardani</surname></string-name>, <string-name><given-names>Guilin</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Aysegul</given-names> <surname>Dundar</surname></string-name>, <string-name><given-names>Shiqiu</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Andrew</given-names> <surname>Tao</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Bryan</given-names> <surname>Catanzaro</surname></string-name></person-group>. <article-title>&#x201C;Neural FFTs for Universal Texture Image Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>33</volume> (<year>2020</year>) (<comment>cit. on pp. 72, 74, 75, 85, 107, 128</comment>).</mixed-citation></ref>
<ref id="CIT264"><label>[MFS08]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Bruce A</given-names> <surname>Maxwell</surname></string-name>, <string-name><given-names>Richard M</given-names> <surname>Friedhoff</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Casey A</given-names> <surname>Smith</surname></string-name></person-group>. <chapter-title>&#x201C;A biilluminant dichromatic reflection model for understanding images&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>2008 IEEE Conference on Computer Vision and Pattern Recognition</italic></source>. <publisher-name>IEEE</publisher-name>. <year>2008</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>8</lpage> (<comment>cit. on p. 10</comment>).</mixed-citation></ref>
<ref id="CIT265"><label>[Maz+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ilya</given-names> <surname>Mazlov</surname></string-name>, <string-name><given-names>Sebastian</given-names> <surname>Merzbach</surname></string-name>, <string-name><given-names>Elena</given-names> <surname>Trunz</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <article-title>&#x201C;Neural Appearance Synthesis and Transfer&#x201D;</article-title>. <comment>In</comment>: (<year>2019</year>), <comment>pp.</comment> <fpage>35</fpage>&#x2013;<lpage>39</lpage> (<comment>cit. on p. 31</comment>).</mixed-citation></ref>
<ref id="CIT266"><label>[MHM18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Leland</given-names> <surname>McInnes</surname></string-name>, <string-name><given-names>John</given-names> <surname>Healy</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>James</given-names> <surname>Melville</surname></string-name></person-group>. <article-title>&#x201C;Umap: Uniform Manifold Approximation and Projection for Dimension Reduction&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1802.03426</italic></source> (<year>2018</year>) (<comment>cit. on p. 133</comment>).</mixed-citation></ref>
<ref id="CIT267"><label>[Meh+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ishit</given-names> <surname>Mehta</surname></string-name>, <string-name><given-names>Micha&#x00EB;l</given-names> <surname>Gharbi</surname></string-name>, <string-name><given-names>Connelly</given-names> <surname>Barnes</surname></string-name>, <string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name>, <string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Manmohan</given-names> <surname>Chandraker</surname></string-name></person-group>. <article-title>&#x201C;Modulated Periodic Activations for Generalizable Local Functional Representations&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>14214</fpage>&#x2013;<lpage>14223</lpage> (<comment>cit. on pp. 109, 186, 203</comment>).</mixed-citation></ref>
<ref id="CIT268"><label>[MR21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sachin</given-names> <surname>Mehta</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mohammad</given-names> <surname>Rastegari</surname></string-name></person-group>. <article-title>&#x201C;MobileViT: Light-Weight, General-Purpose, and Mobile-Friendly Vision Transformer&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2021</year> (<comment>cit. on pp. 124, 127, 132, 141, 144</comment>).</mixed-citation></ref>
<ref id="CIT269"><label>[Mel+12]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Francho</given-names> <surname>Melendez</surname></string-name>, <string-name><given-names>Mashhuda</given-names> <surname>Glencross</surname></string-name>, <string-name><given-names>Jon</given-names> <surname>Starck</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Gregory J</given-names> <surname>Ward</surname></string-name></person-group>. <article-title>&#x201C;Transfer of Albedo and Local Depth Variation to Photo-Textures&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 9th European <target target-type="page" id="pges_247"/>Conference on Visual Media Production</italic></source>. <year>2012</year>, <comment>pp.</comment> <fpage>40</fpage>&#x2013;<lpage>48</lpage> (<comment>cit. on pp. 31, 34</comment>).</mixed-citation></ref>
<ref id="CIT270"><label>[Mel+21]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Joe</given-names> <surname>Mellor</surname></string-name>, <string-name><given-names>Jack</given-names> <surname>Turner</surname></string-name>, <string-name><given-names>Amos</given-names> <surname>Storkey</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Elliot J</given-names> <surname>Crowley</surname></string-name></person-group>. <chapter-title>&#x201C;Neural Architecture Search Without Training&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>7588</fpage>&#x2013;<lpage>7598</lpage> (<comment>cit. on p. 92</comment>).</mixed-citation></ref>
<ref id="CIT271"><label>[Mer+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sebastian</given-names> <surname>Merzbach</surname></string-name>, <string-name><given-names>Max</given-names> <surname>Hermann</surname></string-name>, <string-name><given-names>Martin</given-names> <surname>Rump</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <article-title>&#x201C;Learned Fitting of Spatially Varying BRDFs&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 38</volume>. <issue>4</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>193</fpage>&#x2013;<lpage>205</lpage> (<comment>cit. on p. 30</comment>).</mixed-citation></ref>
<ref id="CIT272"><label>[MWK17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sebastian</given-names> <surname>Merzbach</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Weinmann</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <article-title>&#x201C;High-quality Multi-Spectral Reflectance Acquisition with X-rite TAC7&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the Workshop on Material Appearance Modeling</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>11</fpage>&#x2013;<lpage>16</lpage> (<comment>cit. on pp. 35, 39</comment>).</mixed-citation></ref>
<ref id="CIT273"><label>[MNG17]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Lars</given-names> <surname>Mescheder</surname></string-name>, <string-name><given-names>Sebastian</given-names> <surname>Nowozin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Andreas</given-names> <surname>Geiger</surname></string-name></person-group>. <chapter-title>&#x201C;Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2017</year>, <comment>pp.</comment> <fpage>2391</fpage>&#x2013;<lpage>2400</lpage> (<comment>cit. on p. 126</comment>).</mixed-citation></ref>
<ref id="CIT274"><label>[Mic+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Paulius</given-names> <surname>Micikevicius</surname></string-name>, <string-name><given-names>Sharan</given-names> <surname>Narang</surname></string-name>, <string-name><given-names>Jonah</given-names> <surname>Alben</surname></string-name>, <string-name><given-names>Gregory</given-names> <surname>Diamos</surname></string-name>, <string-name><given-names>Erich</given-names> <surname>Elsen</surname></string-name>, <string-name><given-names>David</given-names> <surname>Garcia</surname></string-name>, <string-name><given-names>Boris</given-names> <surname>Ginsburg</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Houston</surname></string-name>, <string-name><given-names>Oleksii</given-names> <surname>Kuchaiev</surname></string-name>, <string-name><given-names>Ganesh</given-names> <surname>Venkatesh</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;Mixed Precision Training&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2018</year> (<comment>cit. on pp. 56, 67, 84, 94, 110, 142, 176, 194</comment>).</mixed-citation></ref>
<ref id="CIT275"><label>[Mig+12]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Eder</given-names> <surname>Miguel</surname></string-name>, <string-name><given-names>Derek</given-names> <surname>Bradley</surname></string-name>, <string-name><given-names>Bernhard</given-names> <surname>Thomaszewski</surname></string-name>, <string-name><given-names>Bernd</given-names> <surname>Bickel</surname></string-name>, <string-name><given-names>Wojciech</given-names> <surname>Matusik</surname></string-name>, <string-name><given-names>Miguel A</given-names> <surname>Otaduy</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Steve</given-names> <surname>Marschner</surname></string-name></person-group>. <article-title>&#x201C;Data-driven Estimation of Cloth Simulation Models&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 31</volume>. <issue>2pt2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2012</year>, <fpage>pp.</fpage> <fpage>519</fpage>&#x2013;<lpage>528</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT276"><label>[Mil+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ben</given-names> <surname>Mildenhall</surname></string-name>, <string-name><given-names>Pratul P.</given-names> <surname>Srinivasan</surname></string-name>, <string-name><given-names>Matthew</given-names> <surname>Tancik</surname></string-name>, <string-name><given-names>Jonathan T.</given-names> <surname>Barron</surname></string-name>, <string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ren</given-names> <surname>Ng</surname></string-name></person-group>. <article-title>&#x201C;NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <year>2020</year> (<comment>cit. on pp. 186, 187, 207, 212, 262, 264</comment>).</mixed-citation></ref>
<ref id="CIT277"><label>[Min95]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Pier</given-names> <surname>Giorgio Minazio</surname></string-name></person-group>. <article-title>&#x201C;FAST&#x2013;Fabric Assurance by Simple Testing&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Journal of Clothing Science and Technology</italic></source> (<year>1995</year>) (<comment>cit. on pp. 156, 158</comment>).</mixed-citation></ref>
<ref id="CIT278"><label>[Miy+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Takeru</given-names> <surname>Miyato</surname></string-name>, <string-name><given-names>Toshiki</given-names> <surname>Kataoka</surname></string-name>, <string-name><given-names>Masanori</given-names> <surname>Koyama</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yuichi</given-names> <surname>Yoshida</surname></string-name></person-group>. <article-title>&#x201C;Spectral Normalization for Generative Adversarial <target target-type="page" id="pges_248"/>Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations</italic></source>. <year>2018</year> (<comment>cit. on pp. 141, 145</comment>).</mixed-citation></ref>
<ref id="CIT279"><label>[Mor+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alexander</given-names> <surname>Mordvintsev</surname></string-name>, <string-name><given-names>Ettore</given-names> <surname>Randazzo</surname></string-name>, <string-name><given-names>Eyvind</given-names> <surname>Niklasson</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Michael</given-names> <surname>Levin</surname></string-name></person-group>. <article-title>&#x201C;Growing Neural Cellular Automata&#x201D;</article-title>. <comment>In</comment>: <source><italic>Distill</italic></source> (<year>2020</year>). <uri xlink:href="https://distill.pub/2020/growing-ca">https://distill.pub/2020/growing-ca</uri> (<comment>cit. on pp. 15, 75, 257</comment>).</mixed-citation></ref>
<ref id="CIT280"><label>[Mor15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Andrew</given-names> <surname>Morgan</surname></string-name></person-group>. <source><italic>The True Cost</italic></source>. <year>2015</year> (<comment>cit. on p. 2</comment>).</mixed-citation></ref>
<ref id="CIT281"><label>[Mor+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Joep</given-names> <surname>Moritz</surname></string-name>, <string-name><given-names>Stuart</given-names> <surname>James</surname></string-name>, <string-name><given-names>Tom S.F.</given-names> <surname>Haines</surname></string-name>, <string-name><given-names>Tobias</given-names> <surname>Ritschel</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Tim</given-names> <surname>Weyrich</surname></string-name></person-group>. <article-title>&#x201C;Texture Stationarization: Turning Photos into Tileable Textures&#x201D;</article-title>. <comment>In</comment>: <source><italic>Eurographics Symposium on Geometry Processing</italic></source>. <volume>Vol. 36</volume>. <issue>2</issue>. <year>2017</year>, <comment>pp.</comment> <fpage>177</fpage>&#x2013;<lpage>188</lpage> (<comment>cit. on pp. 15, 72&#x2013;74, 88&#x2013;90, 99, 114, 257</comment>).</mixed-citation></ref>
<ref id="CIT282"><label>[MMK03]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Gero</given-names> <surname>M&#x00FC;ller</surname></string-name>, <string-name><given-names>Jan</given-names> <surname>Meseth</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <article-title>&#x201C;Compression and Real-Time Rendering of Measured BTFs Using Local PCA.&#x201D;</article-title> <comment>In</comment>: <source><italic>VMV</italic></source>. <year>2003</year>, <comment>pp.</comment> <fpage>271</fpage>&#x2013;<lpage>279</lpage> (<comment>cit. on p. 103</comment>).</mixed-citation></ref>
<ref id="CIT283"><label>[M&#x00FC;l+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>M&#x00FC;ller</surname></string-name>, <string-name><given-names>Alex</given-names> <surname>Evans</surname></string-name>, <string-name><given-names>Christoph</given-names> <surname>Schied</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Alexander</given-names> <surname>Keller</surname></string-name></person-group>. <article-title>&#x201C;Instant Neural Graphics Primitives with a Multiresolution Hash Encoding&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>41</volume>.<issue>4</issue> (<year>2022</year>), <volume>102</volume>:<fpage>1</fpage>&#x2013;102:<lpage>15</lpage> (<comment>cit. on pp. 203, 212</comment>).</mixed-citation></ref>
<ref id="CIT284"><label>[M&#x00FC;l+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>M&#x00FC;ller</surname></string-name>, <string-name><given-names>Brian</given-names> <surname>McWilliams</surname></string-name>, <string-name><given-names>Fabrice</given-names> <surname>Rousselle</surname></string-name>, <string-name><given-names>Markus</given-names> <surname>Gross</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Nov&#x00E1;k</surname></string-name></person-group>. <article-title>&#x201C;Neural Importance Sampling&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>38</volume>.<issue>5</issue> (<year>2019</year>), <fpage>pp.</fpage> <fpage>1</fpage>&#x2013;<lpage>19</lpage> (<comment>cit. on pp. 116, 186, 191, 196, 197</comment>).</mixed-citation></ref>
<ref id="CIT285"><label>[M&#x00FC;l+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>M&#x00FC;ller</surname></string-name>, <string-name><given-names>Fabrice</given-names> <surname>Rousselle</surname></string-name>, <string-name><given-names>Alexander</given-names> <surname>Keller</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Nov&#x00E1;k</surname></string-name></person-group>. <article-title>&#x201C;Neural Control Variates&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>39</volume>.<issue>6</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>19</lpage> (<comment>cit. on pp. 186, 191</comment>).</mixed-citation></ref>
<ref id="CIT286"><label>[Mun+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jacob</given-names> <surname>Munkberg</surname></string-name>, <string-name><given-names>Jon</given-names> <surname>Hasselgren</surname></string-name>, <string-name><given-names>Tianchang</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>Jun</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>Wenzheng</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Alex</given-names> <surname>Evans</surname></string-name>, <string-name><given-names>Thomas</given-names> <surname>M&#x00FC;ller</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Sanja</given-names> <surname>Fidler</surname></string-name></person-group>. <article-title>&#x201C;Extracting Triangular 3D Models, Materials, and Lighting From Images&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>8280</fpage>&#x2013;<lpage>8290</lpage> (<comment>cit. on pp. 127, 129</comment>).</mixed-citation></ref>
<ref id="CIT287"><label>[Mur+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J Krishna</given-names> <surname>Murthy</surname></string-name>, <string-name><given-names>Miles</given-names> <surname>Macklin</surname></string-name>, <string-name><given-names>Florian</given-names> <surname>Golemo</surname></string-name>, <string-name><given-names>Vikram</given-names> <surname>Voleti</surname></string-name>, <string-name><given-names>Linda</given-names> <surname>Petrini</surname></string-name>, <string-name><given-names>Martin</given-names> <surname>Weiss</surname></string-name>, <string-name><given-names>Breandan</given-names> <surname>Considine</surname></string-name>, <string-name><given-names>J&#x00E9;r&#x00F4;me</given-names> <surname>Parent-L&#x00E9;vesque</surname></string-name>, <string-name><given-names>Kevin</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>Kenny</given-names> <surname>Erleben</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;GradSim: Differentiable simulation for System Identification and Visuomotor Control&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2020</year> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT288"><label>[Naf+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Hossein Ziaei</given-names> <surname>Nafchi</surname></string-name>, <string-name><given-names>Atena</given-names> <surname>Shahkolaei</surname></string-name>, <string-name><given-names>Rachid</given-names> <surname>Hedjam</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mohamed</given-names> <surname>Cheriet</surname></string-name></person-group>. <article-title>&#x201C;Mean Deviation Similarity Index: Efficient <target target-type="page" id="pges_249"/>and Reliable Full-reference Image Quality Evaluator&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Access</italic></source> <volume>4</volume> (<year>2016</year>), <comment>pp.</comment> <fpage>5579</fpage>&#x2013;<lpage>5590</lpage> (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT289"><label>[Nag+15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Koki</given-names> <surname>Nagano</surname></string-name>, <string-name><given-names>Graham</given-names> <surname>Fyffe</surname></string-name>, <string-name><given-names>Oleg</given-names> <surname>Alexander</surname></string-name>, <string-name><given-names>Jernej</given-names> <surname>Barbic</surname></string-name>, <string-name><given-names>Hao</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Abhijeet</given-names> <surname>Ghosh</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Paul E</given-names> <surname>Debevec</surname></string-name></person-group>. <article-title>&#x201C;Skin Microstructure Deformation with Displacement Map Convolution.&#x201D;</article-title> <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>34</volume>.<issue>4</issue> (<year>2015</year>), <comment>pp.</comment> <fpage>109</fpage>&#x2013;<lpage>1</lpage> (<comment>cit. on p. 105</comment>).</mixed-citation></ref>
<ref id="CIT290"><label>[NH10]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Vinod</given-names> <surname>Nair</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Geoffrey E</given-names> <surname>Hinton</surname></string-name></person-group>. <article-title>&#x201C;Rectified Linear Units Improve Restricted Boltzmann Machines&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <year>2010</year> (<comment>cit. on pp. 54, 109, 111, 120</comment>).</mixed-citation></ref>
<ref id="CIT291"><label>[Nam+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Giljoo</given-names> <surname>Nam</surname></string-name>, <string-name><given-names>Joo</given-names> <surname>Ho Lee</surname></string-name>, <string-name><given-names>Hongzhi</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Diego</given-names> <surname>Gutierrez</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Min H</given-names> <surname>Kim</surname></string-name></person-group>. <article-title>&#x201C;Simultaneous Acquisition of Microscale Reflectance and Normals&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>35</volume>.<issue>6</issue> (<year>2016</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on pp. 35, 39</comment>).</mixed-citation></ref>
<ref id="CIT292"><label>[NK18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Hyeonseob</given-names> <surname>Nam</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hyo</given-names> <surname>Eun Kim</surname></string-name></person-group>. <article-title>&#x201C;Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <volume>Vol. 2018-Decem</volume>. <year>2018</year>, <comment>pp.</comment> <fpage>2558</fpage>&#x2013;<lpage>2567</lpage> (<comment>cit. on p. 77</comment>).</mixed-citation></ref>
<ref id="CIT293"><label>[NSO12]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Rahul</given-names> <surname>Narain</surname></string-name>, <string-name><given-names>Armin</given-names> <surname>Samii</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>James F</given-names> <surname>O&#x2019;brien</surname></string-name></person-group>. <article-title>&#x201C;Adaptive Anisotropic Remeshing for Cloth Simulation&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>31</volume>.<issue>6</issue> (<year>2012</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>10</lpage> (<comment>cit. on p. 162</comment>).</mixed-citation></ref>
<ref id="CIT294"><label>[ND06]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Addy</given-names> <surname>Ngan</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Fr&#x00E9;do</given-names> <surname>Durand</surname></string-name></person-group>. <article-title>&#x201C;Statistical Acquisition of Texture Appearance&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 17th Eurographics Conference on Rendering Techniques</italic></source>. <year>2006</year>, <fpage>pp.</fpage> <fpage>31</fpage>&#x2013;<lpage>40</lpage> (<comment>cit. on p. 63</comment>).</mixed-citation></ref>
<ref id="CIT295"><label>[NJR15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jannik</given-names> <surname>Boll Nielsen</surname></string-name>, <string-name><given-names>Henrik</given-names> <surname>Wann Jensen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name></person-group>. <article-title>&#x201C;On Optimal, Minimal BRDF Sampling for Reflectance Acquisition&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>34</volume>.<issue>6</issue> (<year>2015</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on pp. 42, 116, 130</comment>).</mixed-citation></ref>
<ref id="CIT296"><label>[Nik+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Eyvind</given-names> <surname>Niklasson</surname></string-name>, <string-name><given-names>Alexander</given-names> <surname>Mordvintsev</surname></string-name>, <string-name><given-names>Ettore</given-names> <surname>Randazzo</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Michael</given-names> <surname>Levin</surname></string-name></person-group>. <article-title>&#x201C;Self-Organising Textures&#x201D;</article-title>. <comment>In</comment>: <source><italic>Distill</italic></source> (<year>2021</year>). <uri xlink:href="https://distill.pub/selforg/2021/textures">https://distill.pub/selforg/2021/textures</uri> (<comment>cit. on pp. 72, 74, 75, 89&#x2013;91, 93, 95, 99</comment>).</mixed-citation></ref>
<ref id="CIT297"><label>[Nim+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Merlin</given-names> <surname>Nimier-David</surname></string-name>, <string-name><given-names>Delio</given-names> <surname>Vicini</surname></string-name>, <string-name><given-names>Tizian</given-names> <surname>Zeltner</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Wenzel</given-names> <surname>Jakob</surname></string-name></person-group>. <article-title>&#x201C;Mitsuba 2: A Retargetable Forward and Inverse Renderer&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>38</volume>.<issue>6</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>17</lpage> (<comment>cit. on pp. 6, 109, 202</comment>).</mixed-citation></ref>
<ref id="CIT298"><label>[ODO16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Augustus</given-names> <surname>Odena</surname></string-name>, <string-name><given-names>Vincent</given-names> <surname>Dumoulin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Chris</given-names> <surname>Olah</surname></string-name></person-group>. <article-title>&#x201C;Deconvolution and Checker-board Artifacts&#x201D;</article-title>. <comment>In</comment>: <source><italic>Distill</italic></source> (<year>2016</year>) (<comment>cit. on p. 109</comment>).</mixed-citation></ref>
<ref id="CIT299"><label>[Ouy+21]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_250"/><person-group person-group-type="author"><string-name><given-names>Yaobin</given-names> <surname>Ouyang</surname></string-name>, <string-name><given-names>Shiqiu</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Markus</given-names> <surname>Kettunen</surname></string-name>, <string-name><given-names>Matt</given-names> <surname>Pharr</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jacopo</given-names> <surname>Pantaleoni</surname></string-name></person-group>. <article-title>&#x201C;ReSTIR GI: Path Resampling for Real-time Path Tracing&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 40</volume>. <issue>8</issue>. <year>2021</year>, <comment>pp.</comment> <fpage>17</fpage>&#x2013;<lpage>29</lpage> (<comment>cit. on p. 184</comment>).</mixed-citation></ref>
<ref id="CIT300"><label>[Pap+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>George</given-names> <surname>Papamakarios</surname></string-name>, <string-name><given-names>Eric</given-names> <surname>Nalisnick</surname></string-name>, <string-name><given-names>Danilo</given-names> <surname>Jimenez Rezende</surname></string-name>, <string-name><given-names>Shakir</given-names> <surname>Mohamed</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Balaji</given-names> <surname>Lakshminarayanan</surname></string-name></person-group>. <article-title>&#x201C;Normalizing Flows for Probabilistic Modeling and Inference&#x201D;</article-title>. <comment>In</comment>: <source><italic>The Journal of Machine Learning Research</italic></source> <volume>22</volume>.<issue>1</issue> (<year>2021</year>), <comment>pp.</comment> <fpage>2617</fpage>&#x2013;<lpage>2680</lpage> (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT301"><label>[PPM17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>George</given-names> <surname>Papamakarios</surname></string-name>, <string-name><given-names>Theo</given-names> <surname>Pavlakou</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Iain</given-names> <surname>Murray</surname></string-name></person-group>. <article-title>&#x201C;Masked Autoregressive Flow for Density Estimation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>30</volume> (<year>2017</year>) (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT302"><label>[PSM19]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>George</given-names> <surname>Papamakarios</surname></string-name>, <string-name><given-names>David</given-names> <surname>Sterratt</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Iain</given-names> <surname>Murray</surname></string-name></person-group>. <article-title>&#x201C;Sequential Neural Likelihood: Fast Likelihood-free Inference with Autoregressive Flows&#x201D;</article-title>. <comment>In</comment>: <source><italic>The 22nd International Conference on Artificial Intelligence and Statistics</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>837</fpage>&#x2013;<lpage>848</lpage> (<comment>cit. on p. 191</comment>).</mixed-citation></ref>
<ref id="CIT303"><label>[Par+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Taesung</given-names> <surname>Park</surname></string-name>, <string-name><given-names>Alexei A</given-names> <surname>Efros</surname></string-name>, <string-name><given-names>Richard</given-names> <surname>Zhang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jun-Yan</given-names> <surname>Zhu</surname></string-name></person-group>. <article-title>&#x201C;Contrastive Learning for Unpaired Image-to-Image Translation&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2007.15651</italic></source> (<year>2020</year>) (<comment>cit. on pp. 37, 210</comment>).</mixed-citation></ref>
<ref id="CIT304"><label>[Pas+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Adam</given-names> <surname>Paszke</surname></string-name>, <string-name><given-names>Sam</given-names> <surname>Gross</surname></string-name>, <string-name><given-names>Francisco</given-names> <surname>Massa</surname></string-name>, <string-name><given-names>Adam</given-names> <surname>Lerer</surname></string-name>, <string-name><given-names>James</given-names> <surname>Bradbury</surname></string-name>, <string-name><given-names>Gregory</given-names> <surname>Chanan</surname></string-name>, <string-name><given-names>Trevor</given-names> <surname>Killeen</surname></string-name>, <string-name><given-names>Zeming</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Natalia</given-names> <surname>Gimelshein</surname></string-name>, <string-name><given-names>Luca</given-names> <surname>Antiga</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;Pytorch: An Imperative Style, High-performance Deep Learning Library&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <volume>Vol. 32</volume>. <year>2019</year> (<comment>cit. on pp. 39, 55, 67, 84, 94, 110, 142, 176, 194</comment>).</mixed-citation></ref>
<ref id="CIT305"><label>[Pha+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Minh Quang</given-names> <surname>Pham</surname></string-name>, <string-name><given-names>Josep-Maria</given-names> <surname>Crego</surname></string-name>, <string-name><given-names>Fran&#x00E7;ois</given-names> <surname>Yvon</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jean</given-names> <surname>Senellart</surname></string-name></person-group>. <article-title>&#x201C;A study of Residual Adapters for Multi-domain Neural Machine Translation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Conference on Machine Translation</italic></source>. <year>2020</year> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT306"><label>[Pha19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Matt</given-names> <surname>Pharr</surname></string-name></person-group>. <source><italic>Visualizing Warping Strategies for Sampling Environment Map Lights</italic></source>. <year>2019</year> (<comment>cit. on pp. 189, 199, 200</comment>).</mixed-citation></ref>
<ref id="CIT307"><label>[PJH20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Matt</given-names> <surname>Pharr</surname></string-name>, <string-name><given-names>Wenzel</given-names> <surname>Jakob</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Greg</given-names> <surname>Humphreys</surname></string-name></person-group>. <source><italic>Implementation of the forthcoming 4th edition of Physically Based Rendering: From Theory to Implementation</italic></source>. <year>2020</year> (<comment>cit. on pp. 3, 109, 189, 200</comment>).</mixed-citation></ref>
<ref id="CIT308"><label>[PJH16]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Matt</given-names> <surname>Pharr</surname></string-name>, <string-name><given-names>Wenzel</given-names> <surname>Jakob</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Greg</given-names> <surname>Humphreys</surname></string-name></person-group>. <source><italic>Physically Based Rendering: From Theory to Implementation</italic></source>. <edition>3rd</edition>. <publisher-loc>San Francisco, CA, USA</publisher-loc>: <publisher-name>Morgan Kaufmann Publishers Inc.</publisher-name>, <year>2016</year> (<comment>cit. on pp. 185, 189, 200, 201</comment>).</mixed-citation></ref>
<ref id="CIT309"><label>[PS00]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_251"/><person-group person-group-type="author"><string-name><given-names>Javier</given-names> <surname>Portilla</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Eero P</given-names> <surname>Simoncelli</surname></string-name></person-group>. <article-title>&#x201C;A Parametric Texture Model Based on Joint Statistics of Complex Wavelet Coefficients&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Journal of Computer Vision</italic></source> <volume>40</volume>.<issue>1</issue> (<year>2000</year>), <comment>pp.</comment> <fpage>49</fpage>&#x2013;<lpage>70</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT310"><label>[Pow13]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jess</given-names> <surname>Power</surname></string-name></person-group>. <article-title>&#x201C;Fabric Objective Measurements for Commercial 3D Virtual Garment Simulation&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Journal of Clothing Science and Technology</italic></source> (<year>2013</year>) (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT311"><label>[Pra+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ekta</given-names> <surname>Prashnani</surname></string-name>, <string-name><given-names>Hong</given-names> <surname>Cai</surname></string-name>, <string-name><given-names>Yasamin</given-names> <surname>Mostofi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Pradeep</given-names> <surname>Sen</surname></string-name></person-group>. <article-title>&#x201C;Pieapp: Perceptual Image-error Assessment Through Pairwise Preference&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2018</year>, <comment>pp.</comment> <fpage>1808</fpage>&#x2013;<lpage>1817</lpage> (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT312"><label>[Raa+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Lara</given-names> <surname>Raad</surname></string-name>, <string-name><given-names>Axel</given-names> <surname>Davy</surname></string-name>, <string-name><given-names>Agn&#x00E8;s</given-names> <surname>Desolneux</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jean-Michel</given-names> <surname>Morel</surname></string-name></person-group>. <article-title>&#x201C;A Survey of Exemplar-Based Texture Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Annals of Mathematical Sciences and Applications</italic></source> <volume>3</volume>.<issue>1</issue> (<year>2018</year>), <comment>pp.</comment> <fpage>89</fpage>&#x2013;<lpage>148</lpage> (<comment>cit. on p. 73</comment>).</mixed-citation></ref>
<ref id="CIT313"><label>[Rad+21]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Alec</given-names> <surname>Radford</surname></string-name>, <string-name><given-names>Jong</given-names> <surname>Wook Kim</surname></string-name>, <string-name><given-names>Chris</given-names> <surname>Hallacy</surname></string-name>, <string-name><given-names>Aditya</given-names> <surname>Ramesh</surname></string-name>, <string-name><given-names>Gabriel</given-names> <surname>Goh</surname></string-name>, <string-name><given-names>Sandhini</given-names> <surname>Agarwal</surname></string-name>, <string-name><given-names>Girish</given-names> <surname>Sastry</surname></string-name>, <string-name><given-names>Amanda</given-names> <surname>Askell</surname></string-name>, <string-name><given-names>Pamela</given-names> <surname>Mishkin</surname></string-name>, <string-name><given-names>Jack</given-names> <surname>Clark</surname></string-name></person-group>, <etal>et al.</etal> <chapter-title>&#x201C;Learning Transferable Visual Models from Natural Language Supervision&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>8748</fpage>&#x2013;<lpage>8763</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT314"><label>[Rag+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Maithra</given-names> <surname>Raghu</surname></string-name>, <string-name><given-names>Chiyuan</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Jon</given-names> <surname>Kleinberg</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Samy</given-names> <surname>Bengio</surname></string-name></person-group>. <article-title>&#x201C;Transfusion: Under-standing Transfer Learning for Medical Imaging&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>32</volume> (<year>2019</year>) (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT315"><label>[Rai+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Gilles</given-names> <surname>Rainer</surname></string-name>, <string-name><given-names>Adrien</given-names> <surname>Bousseau</surname></string-name>, <string-name><given-names>Tobias</given-names> <surname>Ritschel</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>George</given-names> <surname>Drettakis</surname></string-name></person-group>. <article-title>&#x201C;Neural Pre-computed Radiance Transfer&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 41</volume>. <issue>2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2022</year>, <comment>pp.</comment> <fpage>365</fpage>&#x2013;<lpage>378</lpage> (<comment>cit. on pp. 105, 116, 203</comment>).</mixed-citation></ref>
<ref id="CIT316"><label>[Rai+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Gilles</given-names> <surname>Rainer</surname></string-name>, <string-name><given-names>Abhijeet</given-names> <surname>Ghosh</surname></string-name>, <string-name><given-names>Wenzel</given-names> <surname>Jakob</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Tim</given-names> <surname>Weyrich</surname></string-name></person-group>. <article-title>&#x201C;Unified Neural Encod-ing of BTFs&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 39</volume>. <issue>2</issue>. <publisher-name>The Eurographics Association</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>13</lpage> (<comment>cit. on pp. 18, 33, 102, 103, 109, 112, 113, 116, 259</comment>).</mixed-citation></ref>
<ref id="CIT317"><label>[Rai+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Gilles</given-names> <surname>Rainer</surname></string-name>, <string-name><given-names>Wenzel</given-names> <surname>Jakob</surname></string-name>, <string-name><given-names>Abhijeet</given-names> <surname>Ghosh</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Tim</given-names> <surname>Weyrich</surname></string-name></person-group>. <article-title>&#x201C;Neural BTF Compres-sion and Interpolation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 38</volume>. <issue>2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>235</fpage>&#x2013;<lpage>244</lpage> (<comment>cit. on pp. 18, 33, 39, 40, 102, 103, 106, 109, 112, 113, 259</comment>).</mixed-citation></ref>
<ref id="CIT318"><label>[Ram+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Prajit</given-names> <surname>Ramachandran</surname></string-name>, <string-name><given-names>Niki</given-names> <surname>Parmar</surname></string-name>, <string-name><given-names>Ashish</given-names> <surname>Vaswani</surname></string-name>, <string-name><given-names>Irwan</given-names> <surname>Bello</surname></string-name>, <string-name><given-names>Anselm</given-names> <surname>Levskaya</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jon</given-names> <surname>Shlens</surname></string-name></person-group>. <article-title>&#x201C;Stand-alone Self-attention <target target-type="page" id="pges_252"/>in Vision Models&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>32</volume> (<year>2019</year>) (<comment>cit. on p. 164</comment>).</mixed-citation></ref>
<ref id="CIT319"><label>[RH01a]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Pat</given-names> <surname>Hanrahan</surname></string-name></person-group>. <chapter-title>&#x201C;An Efficient Representation for Irradiance Environment Maps&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques</italic></source>. <comment>SIGGRAPH &#x2019;01</comment>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>, <year>2001</year>, <comment>pp.</comment> <fpage>497</fpage>&#x2013;<lpage>500</lpage> (<comment>cit. on p. 13</comment>).</mixed-citation></ref>
<ref id="CIT320"><label>[RH01b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Pat</given-names> <surname>Hanrahan</surname></string-name></person-group>. <article-title>&#x201C;An Efficient Representation for Irradiance Environment Maps&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques</italic></source>. <year>2001</year>, <comment>pp.</comment> <fpage>497</fpage>&#x2013;<lpage>500</lpage> (<comment>cit. on pp. 196, 199, 202</comment>).</mixed-citation></ref>
<ref id="CIT321"><label>[Ras+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Abdullah</given-names> <surname>Haroon Rasheed</surname></string-name>, <string-name><given-names>Victor</given-names> <surname>Romero</surname></string-name>, <string-name><given-names>Florence</given-names> <surname>Bertails-Descoubes</surname></string-name>, <string-name><given-names>Stefanie</given-names> <surname>Wuhrer</surname></string-name>, <string-name><given-names>Jean-S&#x00E9;bastien</given-names> <surname>Franco</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Arnaud</given-names> <surname>Lazarus</surname></string-name></person-group>. <article-title>&#x201C;Learning to Measure the Static Friction Coefficient in Cloth Contact&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>9912</fpage>&#x2013;<lpage>9921</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT322"><label>[RD17]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Ajita</given-names> <surname>Rattani</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reza</given-names> <surname>Derakhshani</surname></string-name></person-group>. <chapter-title>&#x201C;On Fine-tuning Convolutional Neural Networks for Smartphone Based Ocular Recognition&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>2017 IEEE international joint conference on biometrics (IJCB)</italic></source>. <publisher-name>IEEE</publisher-name>. <year>2017</year>, <comment>pp.</comment> <fpage>762</fpage>&#x2013;<lpage>767</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT323"><label>[RBV18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sylvestre-Alvise</given-names> <surname>Rebuffi</surname></string-name>, <string-name><given-names>Hakan</given-names> <surname>Bilen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Andrea</given-names> <surname>Vedaldi</surname></string-name></person-group>. <article-title>&#x201C;Efficient Parametrization of Multi-domain Deep Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2018</year>, <comment>pp.</comment> <fpage>8119</fpage>&#x2013;<lpage>8127</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT324"><label>[RBV17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sylvestre-Alvise</given-names> <surname>Rebuffi</surname></string-name>, <string-name><given-names>Hakan</given-names> <surname>Bilen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Andrea</given-names> <surname>Vedaldi</surname></string-name></person-group>. <article-title>&#x201C;Learning Multiple Visual Domains with Residual Adapters&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>30</volume> (<year>2017</year>) (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT325"><label>[RJ19]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>A Sai</given-names> <surname>Bharadwaj Reddy</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>D Sujitha</given-names> <surname>Juliet</surname></string-name></person-group>. <chapter-title>&#x201C;Transfer Learning with ResNet-50 for Malaria Cell-image Classification&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE International Conference on Computer Vision (ICCV)</italic></source>. <publisher-name>IEEE</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>0945</fpage>&#x2013;<lpage>0949</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT326"><label>[Ree+15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Scott E</given-names> <surname>Reed</surname></string-name>, <string-name><given-names>Yi</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Yuting</given-names> <surname>Zhang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Honglak</given-names> <surname>Lee</surname></string-name></person-group>. <article-title>&#x201C;Deep Visual Analogy-Making&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2015</year>, <comment>pp.</comment> <fpage>1252</fpage>&#x2013;<lpage>1260</lpage> (<comment>cit. on p. 32</comment>).</mixed-citation></ref>
<ref id="CIT327"><label>[Rei+14]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Klein</given-names> <surname>Reinhard</surname></string-name>, <string-name><given-names>Rump</given-names> <surname>Martin</surname></string-name>, <string-name><given-names>Weinmann</given-names> <surname>Michael</surname></string-name>, <string-name><given-names>Sarlette</given-names> <surname>Ralf</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Schwartz</given-names> <surname>Christopher</surname></string-name></person-group>. <article-title>&#x201C;Design and <target target-type="page" id="pges_253"/>Implementation of Practical Bidirectional Texture Function Measurement Devices focusing on the Developments at the University of Bonn&#x201D;</article-title>. <comment>In</comment>: <source><italic>Sensors</italic></source> <volume>14</volume>.<month>5</month> (<year>2014</year>) (<comment>cit. on pp. 109, 114, 117</comment>).</mixed-citation></ref>
<ref id="CIT328"><label>[Rei+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Rafael</given-names> <surname>Reisenhofer</surname></string-name>, <string-name><given-names>Sebastian</given-names> <surname>Bosse</surname></string-name>, <string-name><given-names>Gitta</given-names> <surname>Kutyniok</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>Wiegand</surname></string-name></person-group>. <article-title>&#x201C;A Haar Wavelet-based Perceptual Similarity Index for Image Quality Assessment&#x201D;</article-title>. <comment>In</comment>: <source><italic>Signal Processing: Image Communication</italic></source> <volume>61</volume> (<year>2018</year>), <comment>pp.</comment> <fpage>33</fpage>&#x2013;<lpage>43</lpage> (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT329"><label>[RSS16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Nathalie</given-names> <surname>Remy</surname></string-name>, <string-name><given-names>Eveline</given-names> <surname>Speelman</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Steven</given-names> <surname>Swartz</surname></string-name></person-group>. <source><italic>Style that&#x2019;s Sustainable: A New Fast-Fashion Formula</italic></source>. <year>2016</year> (<comment>cit. on p. 2</comment>).</mixed-citation></ref>
<ref id="CIT330"><label>[Ren+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jian</given-names> <surname>Ren</surname></string-name>, <string-name><given-names>Xiaohui</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>Zhe</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Radomir</given-names> <surname>Mech</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>David J</given-names> <surname>Foran</surname></string-name></person-group>. <article-title>&#x201C;Personalized Image Aesthetics&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>638</fpage>&#x2013;<lpage>647</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT331"><label>[Rez+20]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Danilo</given-names> <surname>Jimenez Rezende</surname></string-name>, <string-name><given-names>George</given-names> <surname>Papamakarios</surname></string-name>, <string-name><given-names>S&#x00E9;bastien</given-names> <surname>Racaniere</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Albergo</surname></string-name>, <string-name><given-names>Gurtej</given-names> <surname>Kanwar</surname></string-name>, <string-name><given-names>Phiala</given-names> <surname>Shanahan</surname></string-name></person-group>, <comment>and</comment> <string-name><given-names>Kyle</given-names> <surname>Cranmer</surname></string-name>. <chapter-title>&#x201C;Normalizing Flows on Tori and Spheres&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>8083</fpage>&#x2013;<lpage>8092</lpage> (<comment>cit. on p. 189</comment>).</mixed-citation></ref>
<ref id="CIT332"><label>[Al-+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Rami</given-names> <surname>Al-Rfou</surname></string-name>, <string-name><given-names>Guillaume</given-names> <surname>Alain</surname></string-name>, <string-name><given-names>Amjad</given-names> <surname>Almahairi</surname></string-name>, <string-name><given-names>Christof</given-names> <surname>Angermueller</surname></string-name>, <string-name><given-names>Dzmitry</given-names> <surname>Bahdanau</surname></string-name>, <string-name><given-names>Nicolas</given-names> <surname>Ballas</surname></string-name>, <string-name><given-names>Fr&#x00E9;d&#x00E9;ric</given-names> <surname>Bastien</surname></string-name>, <string-name><given-names>Justin</given-names> <surname>Bayer</surname></string-name>, <string-name><given-names>Anatoly</given-names> <surname>Belikov</surname></string-name>, <string-name><given-names>Alexander</given-names> <surname>Belopolsky</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;Theano: A Python Framework for Fast Computation of Mathematical Expressions&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv e-prints</italic></source> (<year>2016</year>), <comment>arXiv&#x2013;1605</comment> (<comment>cit. on p. 95</comment>).</mixed-citation></ref>
<ref id="CIT333"><label>[Rho+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Rho</surname></string-name>, <string-name><given-names>Junwoo</given-names> <surname>Cho</surname></string-name>, <string-name><given-names>Jong</given-names> <surname>Hwan Ko</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Eunbyung</given-names> <surname>Park</surname></string-name></person-group>. <article-title>&#x201C;Neural Residual Flow Fields for Efficient Video Representations&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the Asian Conference on Computer Vision</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>3447</fpage>&#x2013;<lpage>3463</lpage> (<comment>cit. on p. 187</comment>).</mixed-citation></ref>
<ref id="CIT334"><label>[Rib+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Edgar</given-names> <surname>Riba</surname></string-name>, <string-name><given-names>Dmytro</given-names> <surname>Mishkin</surname></string-name>, <string-name><given-names>Daniel</given-names> <surname>Ponsa</surname></string-name>, <string-name><given-names>Ethan</given-names> <surname>Rublee</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Gary</given-names> <surname>Bradski</surname></string-name></person-group>. <article-title>&#x201C;Kornia: An Open Source Differentiable Computer Vision Library for Pytorch&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>3674</fpage>&#x2013;<lpage>3683</lpage> (<comment>cit. on pp. 67, 111, 142, 176</comment>).</mixed-citation></ref>
<ref id="CIT335"><label>[Ric+16]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Stephan R.</given-names> <surname>Richter</surname></string-name>, <string-name><given-names>Vibhav</given-names> <surname>Vineet</surname></string-name>, <string-name><given-names>Stefan</given-names> <surname>Roth</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Vladlen</given-names> <surname>Koltun</surname></string-name></person-group>. <chapter-title>&#x201C;Playing for Data: Ground Truth from Computer Games&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>European Conference on Computer Vision (ECCV)</italic></source>. <role>Ed.</role> <comment>by</comment> <person-group person-group-type="editor"><string-name><given-names>Bastian</given-names> <surname>Leibe</surname></string-name>, <string-name><given-names>Jiri</given-names> <surname>Matas</surname></string-name>, <string-name><given-names>Nicu</given-names> <surname>Sebe</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="editor"><string-name><given-names>Max</given-names> <surname>Welling</surname></string-name></person-group>. <volume>Vol. 9906</volume>. <publisher-name>LNCS. Springer International Publishing</publisher-name>, <year>2016</year>, <comment>pp.</comment> <fpage>102</fpage>&#x2013;<lpage>118</lpage> (<comment>cit. on p. 212</comment>).</mixed-citation></ref>
<ref id="CIT336"><label>[RPG16]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_254"/><person-group person-group-type="author"><string-name><given-names>J&#x00E9;r&#x00E9;my</given-names> <surname>Riviere</surname></string-name>, <string-name><given-names>Pieter</given-names> <surname>Peers</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Abhijeet</given-names> <surname>Ghosh</surname></string-name></person-group>. <article-title>&#x201C;Mobile Surface Reflectometry&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 35</volume>. <issue>1</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2016</year>, <comment>pp.</comment> <fpage>191</fpage>&#x2013;<lpage>202</lpage> (<comment>cit. on pp. 31, 34</comment>).</mixed-citation></ref>
<ref id="CIT337"><label>[Rod+23a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Rodriguez-Pardo</surname></string-name>, <string-name><given-names>Henar</given-names> <surname>Dominguez-Elvira</surname></string-name>, <string-name><given-names>David</given-names> <surname>Pascual-Hernandez</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name></person-group>. <article-title>&#x201C;UMat: Uncertainty-Aware Single Image High Resolution Material Capture&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source> (<year>2023</year>) (<comment>cit. on pp. 24, 62, 76, 102, 107, 108, 110, 116, 193, 268</comment>).</mixed-citation></ref>
<ref id="CIT338"><label>[RG21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Rodriguez-Pardo</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name></person-group>. <article-title>&#x201C;Neural Photometry-guided Visual Attribute Transfer&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Visualization and Computer Graphics</italic></source> (<year>2021</year>) (<comment>cit. on pp. 23, 25, 63&#x2013;70, 102, 105, 107, 108, 110, 114, 125, 129, 131, 142, 193, 268</comment>).</mixed-citation></ref>
<ref id="CIT339"><label>[RG22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Rodriguez-Pardo</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name></person-group>. <article-title>&#x201C;SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Visualization and Computer Graphics</italic></source> (<year>2022</year>) (<comment>cit. on pp. 23, 25, 64, 67, 69, 104, 107, 108, 110, 114, 127, 128, 141, 268</comment>).</mixed-citation></ref>
<ref id="CIT340"><label>[Rod+23b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Rodriguez-Pardo</surname></string-name>, <string-name><given-names>Melania</given-names> <surname>Prieto-Martin</surname></string-name>, <string-name><given-names>Dan</given-names> <surname>Casas</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name></person-group>. <article-title>&#x201C;How Will It Drape Like? Capturing Fabric Mechanics from Depth Images&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum (Proc. of Eurographics)</italic></source> (<year>2023</year>) (<comment>cit. on pp. 18, 23, 212, 268</comment>).</mixed-citation></ref>
<ref id="CIT341"><label>[Rod+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Rodriguez-Pardo</surname></string-name>, <string-name><given-names>Sergio</given-names> <surname>Suja</surname></string-name>, <string-name><given-names>David</given-names> <surname>Pascual</surname></string-name>, <string-name><given-names>Jorge</given-names> <surname>Lopez-Moreno</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Elena</given-names> <surname>Garces</surname></string-name></person-group>. <article-title>&#x201C;Automatic Extraction and Synthesis of Regular Repeatable Patterns&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computers &#x0026; Graphics</italic></source> <volume>83</volume> (<year>2019</year>), <comment>pp.</comment> <fpage>33</fpage>&#x2013;<lpage>41</lpage> (<comment>cit. on pp. 15, 34, 72, 75, 79, 83, 87, 89&#x2013;91, 93, 95, 99, 114, 257</comment>).</mixed-citation></ref>
<ref id="CIT342"><label>[RB19]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Carlos</given-names> <surname>Rodr&#x0131;guez-Pardo</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hakan</given-names> <surname>Bilen</surname></string-name></person-group>. <article-title>&#x201C;Personalised Aesthetics with Residual Adapters&#x201D;</article-title>. <comment>In</comment>: <source><italic>Iberian Conference on Pattern Recognition and Image Analysis</italic></source>. <publisher-name>Springer</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>508</fpage>&#x2013;<lpage>520</lpage> (<comment>cit. on pp. 74, 90, 159</comment>).</mixed-citation></ref>
<ref id="CIT343"><label>[Rom+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Robin</given-names> <surname>Rombach</surname></string-name>, <string-name><given-names>Andreas</given-names> <surname>Blattmann</surname></string-name>, <string-name><given-names>Dominik</given-names> <surname>Lorenz</surname></string-name>, <string-name><given-names>Patrick</given-names> <surname>Esser</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Bj&#x00F6;rn</given-names> <surname>Ommer</surname></string-name></person-group>. <article-title>&#x201C;High-Resolution Image Synthesis with Latent Diffusion Models&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>10684</fpage>&#x2013;<lpage>10695</lpage> (<comment>cit. on pp. 127, 209, 210</comment>).</mixed-citation></ref>
<ref id="CIT344"><label>[RFB15]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Olaf</given-names> <surname>Ronneberger</surname></string-name>, <string-name><given-names>Philipp</given-names> <surname>Fischer</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Thomas</given-names> <surname>Brox</surname></string-name></person-group>. <chapter-title>&#x201C;U-Net: Convolutional Networks for Biomedical Image Segmentation&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>International Conference on Medical image computing and computer-assisted intervention</italic></source>. <publisher-name>Springer</publisher-name>. <year>2015</year>, <comment>pp.</comment> <fpage>234</fpage>&#x2013;<lpage>241</lpage> (<comment>cit. on pp. 37&#x2013;39, 54, 55, 64, 66, 67, 69, 78, 108, 111, 124, 127, 132, 141</comment>).</mixed-citation></ref>
<ref id="CIT345"><label>[RSK13]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_255"/><person-group person-group-type="author"><string-name><given-names>Roland</given-names> <surname>Ruiters</surname></string-name>, <string-name><given-names>Christopher</given-names> <surname>Schwartz</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <article-title>&#x201C;Example-based Interpolation and Synthesis of Bidirectional Texture Functions&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 32</volume>. <issue>2pt3</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2013</year>, <comment>pp.</comment> <fpage>361</fpage>&#x2013;<lpage>370</lpage> (<comment>cit. on p. 63</comment>).</mixed-citation></ref>
<ref id="CIT346"><label>[RSK10]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Martin</given-names> <surname>Rump</surname></string-name>, <string-name><given-names>Ralf</given-names> <surname>Sarlette</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <chapter-title>&#x201C;Groundtruth Data for Multispectral Bidirectional Texture Functions&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Conference on Colour in Graphics, Imaging, and Vision</italic></source>. <volume>Vol. 2010</volume>. <issue>1</issue>. <publisher-name>Society for Imaging Science and Technology</publisher-name>. <year>2010</year>, <comment>pp.</comment> <fpage>326</fpage>&#x2013;<lpage>331</lpage> (<comment>cit. on p. 116</comment>).</mixed-citation></ref>
<ref id="CIT347"><label>[Run+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tom FH</given-names> <surname>Runia</surname></string-name>, <string-name><given-names>Kirill</given-names> <surname>Gavrilyuk</surname></string-name>, <string-name><given-names>Cees GM</given-names> <surname>Snoek</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Arnold WM</given-names> <surname>Smeulders</surname></string-name></person-group>. <article-title>&#x201C;Cloth in the Wind: A Case Study of Physical Measurement Through Simulation&#x201D;</article-title>. <comment>In</comment>: <target target-type="page" id="pges_256"/><source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>10498</fpage>&#x2013;<lpage>10507</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT348"><label>[Rus+04]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Daniel B</given-names> <surname>Russakoff</surname></string-name>, <string-name><given-names>Carlo</given-names> <surname>Tomasi</surname></string-name>, <string-name><given-names>Torsten</given-names> <surname>Rohlfing</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Calvin R</given-names> <surname>Maurer</surname></string-name></person-group>. <chapter-title>&#x201C;Image Similarity Using Mutual Information of Regions&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2004</year>, <comment>pp.</comment> <fpage>596</fpage>&#x2013;<lpage>607</lpage> (<comment>cit. on p. 132</comment>).</mixed-citation></ref>
<ref id="CIT349"><label>[Sah+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Chitwan</given-names> <surname>Saharia</surname></string-name>, <string-name><given-names>William</given-names> <surname>Chan</surname></string-name>, <string-name><given-names>Huiwen</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>Chris</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>Jonathan</given-names> <surname>Ho</surname></string-name>, <string-name><given-names>Tim</given-names> <surname>Salimans</surname></string-name>, <string-name><given-names>David</given-names> <surname>Fleet</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mohammad</given-names> <surname>Norouzi</surname></string-name></person-group>. <article-title>&#x201C;Palette: Image-to-Image Diffusion Models&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM SIGGRAPH 2022 Conference Proceedings</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>10</lpage> (<comment>cit. on pp. 127, 209</comment>).</mixed-citation></ref>
<ref id="CIT350"><label>[San+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Veit</given-names> <surname>Sandfort</surname></string-name>, <string-name><given-names>Ke</given-names> <surname>Yan</surname></string-name>, Perry <string-name><given-names>J</given-names> <surname>Pickhardt</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ronald M</given-names> <surname>Summers</surname></string-name></person-group>. <article-title>&#x201C;Data Augmentation Using Generative Adversarial Networks (CycleGAN) to Improve Generalizability in CT Segmentation Tasks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Scientific reports</italic></source> <volume>9</volume>.<issue>1</issue> (<year>2019</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>9</lpage> (<comment>cit. on p. 45</comment>).</mixed-citation></ref>
<ref id="CIT351"><label>[SC20]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Shen</given-names> <surname>Sang</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Manmohan</given-names> <surname>Chandraker</surname></string-name></person-group>. <chapter-title>&#x201C;Single-Shot Neural Relighting and SVBRDF Estimation&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2020</year>, <comment>pp.</comment> <fpage>85</fpage>&#x2013;<lpage>101</lpage> (<comment>cit. on p. 125</comment>).</mixed-citation></ref>
<ref id="CIT352"><label>[SOC22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Igor</given-names> <surname>Santesteban</surname></string-name>, <string-name><given-names>Miguel A</given-names> <surname>Otaduy</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Dan</given-names> <surname>Casas</surname></string-name></person-group>. <article-title>&#x201C;SNUG: Self-Supervised Neural Dynamic Garments&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source> (<year>2022</year>) (<comment>cit. on pp. 17, 259</comment>).</mixed-citation></ref>
<ref id="CIT353"><label>[SSK03]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Mirko</given-names> <surname>Sattler</surname></string-name>, <string-name><given-names>Ralf</given-names> <surname>Sarlette</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <article-title>&#x201C;Efficient and Realistic Visualization of Cloth&#x201D;</article-title>. <comment>In</comment>: <source><italic>Rendering techniques</italic></source>. <year>2003</year>, <comment>pp.</comment> <fpage>167</fpage>&#x2013;<lpage>178</lpage> (<comment>cit. on p. 109</comment>).</mixed-citation></ref>
<ref id="CIT354"><label>[SSK20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Edgar</given-names> <surname>Schonfeld</surname></string-name>, <string-name><given-names>Bernt</given-names> <surname>Schiele</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Anna</given-names> <surname>Khoreva</surname></string-name></person-group>. <article-title>&#x201C;A U-Net Based Discriminator for Generative Adversarial Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>8207</fpage>&#x2013;<lpage>8216</lpage> (<comment>cit. on pp. 83, 93, 124, 128, 132, 141, 142</comment>).</mixed-citation></ref>
<ref id="CIT355"><label>[Sch+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Vincent</given-names> <surname>Sch&#x00FC;ssler</surname></string-name>, <string-name><given-names>Eric</given-names> <surname>Heitz</surname></string-name>, <string-name><given-names>Johannes</given-names> <surname>Hanika</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Carsten</given-names> <surname>Dachsbacher</surname></string-name></person-group>. <article-title>&#x201C;Microfacet-based Normal Mapping for Robust Monte Carlo Path Tracing&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>36</volume>.<issue>6</issue> (<year>2017</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>12</lpage> (<comment>cit. on p. 201</comment>).</mixed-citation></ref>
<ref id="CIT356"><label>[See66]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Robert T</given-names> <surname>Seeley</surname></string-name></person-group>. <article-title>&#x201C;Spherical Harmonics&#x201D;</article-title>. <comment>In</comment>: <source><italic>The American Mathematical Monthly</italic></source> <volume>73</volume>.<issue>4P2</issue> (<year>1966</year>), <comment>pp.</comment> <fpage>115</fpage>&#x2013;<lpage>121</lpage> (<comment>cit. on pp. 185, 196</comment>).</mixed-citation></ref>
<ref id="CIT357"><label>[Sel+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Raghavendra</given-names> <surname>Selvan</surname></string-name>, <string-name><given-names>Frederik</given-names> <surname>Faye</surname></string-name>, <string-name><given-names>Jon</given-names> <surname>Middleton</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Akshay</given-names> <surname>Pai</surname></string-name></person-group>. <article-title>&#x201C;Uncertainty Quantification in Medical Image Segmentation with Normalizing Flows&#x201D;</article-title>. <comment>In</comment>: <source><italic>Machine Learning in Medical Imaging</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>80</fpage>&#x2013;<lpage>90</lpage> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT358"><label>[SKK18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Murat</given-names> <surname>Sensoy</surname></string-name>, <string-name><given-names>Lance</given-names> <surname>Kaplan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Melih</given-names> <surname>Kandemir</surname></string-name></person-group>. <article-title>&#x201C;Evidential Deep Learning to Quantify Classification Uncertainty&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>31</volume> (<year>2018</year>) (<comment>cit. on pp. 126, 130</comment>).</mixed-citation></ref>
<ref id="CIT359"><label>[SDM19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tamar</given-names> <surname>Rott Shaham</surname></string-name>, <string-name><given-names>Tali</given-names> <surname>Dekel</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Tomer</given-names> <surname>Michaeli</surname></string-name></person-group>. <article-title>&#x201C;SinGAN: Learning a Generative Model From a Single Natural Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <month>Oct</month>. <year>2019</year> (<comment>cit. on pp. 32, 38, 57, 75, 85, 86, 90</comment>).</mixed-citation></ref>
<ref id="CIT360"><label>[Shi+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Liang</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>Beichen</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name>, <string-name><given-names>Kalyan</given-names> <surname>Sunkavalli</surname></string-name>, <string-name><given-names>Tamy</given-names> <surname>Boubekeur</surname></string-name>, <string-name><given-names>Radomir</given-names> <surname>Mech</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Wojciech</given-names> <surname>Matusik</surname></string-name></person-group>. <article-title>&#x201C;Match: Differentiable Material Graphs for Procedural Material Capture&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> (<year>2020</year>) (<comment>cit. on pp. 76, 125, 137, 138, 143, 151&#x2013;153</comment>).</mixed-citation></ref>
<ref id="CIT361"><label>[Sho+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Assaf</given-names> <surname>Shocher</surname></string-name>, <string-name><given-names>Shai</given-names> <surname>Bagon</surname></string-name>, <string-name><given-names>Phillip</given-names> <surname>Isola</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Michal</given-names> <surname>Irani</surname></string-name></person-group>. <article-title>&#x201C;InGAN: Capturing and Retargeting the "DNA" of a Natural Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <month>Oct</month>. <year>2019</year> (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT362"><label>[SCI18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Assaf</given-names> <surname>Shocher</surname></string-name>, <string-name><given-names>Nadav</given-names> <surname>Cohen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Michal</given-names> <surname>Irani</surname></string-name></person-group>. <article-title>&#x201C;&#x201C;Zero-Shot&#x201D; Super-Resolution Using Deep Internal Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <month>June</month> <year>2018</year> (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT363"><label>[SK19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Connor</given-names> <surname>Shorten</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Taghi M</given-names> <surname>Khoshgoftaar</surname></string-name></person-group>. <article-title>&#x201C;A Survey on Image Data Augmentation for Deep Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>Journal of Big Data</italic></source> <volume>6</volume>.<issue>1</issue> (<year>2019</year>), <comment>p.</comment> <fpage>60</fpage> (<comment>cit. on pp. 36, 45</comment>).</mixed-citation></ref>
<ref id="CIT364"><label>[SZ15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Karen</given-names> <surname>Simonyan</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Andrew</given-names> <surname>Zisserman</surname></string-name></person-group>. <article-title>&#x201C;Very Deep Convolutional Networks for Large-Scale Image Recognition&#x201D;</article-title>. <target target-type="page" id="pges_257"/><comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2015</year> (<comment>cit. on pp. 32, 164</comment>).</mixed-citation></ref>
<ref id="CIT365"><label>[Sin+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Abhishek</given-names> <surname>Sinha</surname></string-name>, <string-name><given-names>Kumar</given-names> <surname>Ayush</surname></string-name>, <string-name><given-names>Jiaming</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Burak</given-names> <surname>Uzkent</surname></string-name>, <string-name><given-names>Hongxia</given-names> <surname>Jin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Stefano</given-names> <surname>Ermon</surname></string-name></person-group>. <article-title>&#x201C;Negative Data Augmentation&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2102.05113</italic></source> (<year>2021</year>) (<comment>cit. on p. 93</comment>).</mixed-citation></ref>
<ref id="CIT366"><label>[Sit+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Vincent</given-names> <surname>Sitzmann</surname></string-name>, <string-name><given-names>Julien</given-names> <surname>Martel</surname></string-name>, <string-name><given-names>Alexander</given-names> <surname>Bergman</surname></string-name>, <string-name><given-names>David</given-names> <surname>Lindell</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Gordon</given-names> <surname>Wetzstein</surname></string-name></person-group>. <article-title>&#x201C;Implicit Neural Representations with Periodic Activation Functions&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>33</volume> (<year>2020</year>) (<comment>cit. on pp. 75, 102, 109, 112, 120, 186, 193, 194, 262</comment>).</mixed-citation></ref>
<ref id="CIT367"><label>[Sit+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Vincent</given-names> <surname>Sitzmann</surname></string-name>, <string-name><given-names>Semon</given-names> <surname>Rezchikov</surname></string-name>, <string-name><given-names>William T.</given-names> <surname>Freeman</surname></string-name>, <string-name><given-names>Joshua B.</given-names> <surname>Tenenbaum</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Fredo</given-names> <surname>Durand</surname></string-name></person-group>. <article-title>&#x201C;Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2021</year> (<comment>cit. on p. 187</comment>).</mixed-citation></ref>
<ref id="CIT368"><label>[SKS02]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Peter-Pike</given-names> <surname>Sloan</surname></string-name>, <string-name><given-names>Jan</given-names> <surname>Kautz</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>John</given-names> <surname>Snyder</surname></string-name></person-group>. <article-title>&#x201C;Precomputed Radiance Transfer for Real-Time Rendering in Dynamic, Low-Frequency Lighting Environments&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>21</volume>.<issue>3</issue> (<month>July</month> <year>2002</year>), <comment>pp.</comment> <fpage>527</fpage>&#x2013;<lpage>536</lpage> (<comment>cit. on p. 13</comment>).</mixed-citation></ref>
<ref id="CIT369"><label>[Sne17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xavier</given-names> <surname>Snelgrove</surname></string-name></person-group>. <article-title>&#x201C;High-resolution Multi-scale Neural Texture Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>SIG-GRAPH Asia 2017 Technical Briefs</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>4</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT370"><label>[Sol+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ava P</given-names> <surname>Soleimany</surname></string-name>, <string-name><given-names>Alexander</given-names> <surname>Amini</surname></string-name>, <string-name><given-names>Samuel</given-names> <surname>Goldman</surname></string-name>, <string-name><given-names>Daniela</given-names> <surname>Rus</surname></string-name>, <string-name><given-names>Sangeeta N</given-names> <surname>Bhatia</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Connor W</given-names> <surname>Coley</surname></string-name></person-group>. <article-title>&#x201C;Evidential Deep Learning for Guided Molecular Property Prediction and Discovery&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACS Central Science</italic></source> <volume>7</volume>.<issue>8</issue> (<year>2021</year>), <comment>pp.</comment> <fpage>1356</fpage>&#x2013;<lpage>1367</lpage> (<comment>cit. on pp. 126, 135</comment>).</mixed-citation></ref>
<ref id="CIT371"><label>[Spe+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Georg</given-names> <surname>Sperl</surname></string-name>, <string-name><given-names>Rosa M</given-names> <surname>S&#x00E1;nchez-Banderas</surname></string-name>, <string-name><given-names>Manwen</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Chris</given-names> <surname>Wojtan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Miguel A</given-names> <surname>Otaduy</surname></string-name></person-group>. <article-title>&#x201C;Estimation of Yarn-level Simulation Models for Production Fabrics&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>41</volume>.<issue>4</issue> (<year>2022</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>15</lpage> (<comment>cit. on p. 163</comment>).</mixed-citation></ref>
<ref id="CIT372"><label>[Sri+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Pratul P</given-names> <surname>Srinivasan</surname></string-name>, <string-name><given-names>Ben</given-names> <surname>Mildenhall</surname></string-name>, <string-name><given-names>Matthew</given-names> <surname>Tancik</surname></string-name>, <string-name><given-names>Jonathan T</given-names> <surname>Barron</surname></string-name>, <string-name><given-names>Richard</given-names> <surname>Tucker</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Noah</given-names> <surname>Snavely</surname></string-name></person-group>. <article-title>&#x201C;Lighthouse: Predicting Lighting Volumes for Spatially-Coherent Illumination&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>8080</fpage>&#x2013;<lpage>8089</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT373"><label>[Sri+14]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Nitish</given-names> <surname>Srivastava</surname></string-name>, <string-name><given-names>Geoffrey</given-names> <surname>Hinton</surname></string-name>, <string-name><given-names>Alex</given-names> <surname>Krizhevsky</surname></string-name>, <string-name><given-names>Ilya</given-names> <surname>Sutskever</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ruslan</given-names> <surname>Salakhutdinov</surname></string-name></person-group>. <article-title>&#x201C;Dropout: A Simple Way <target target-type="page" id="pges_258"/>to Prevent Neural Networks from Overfitting&#x201D;</article-title>. <comment>In</comment>: <source><italic>The Journal of Machine Learning Research</italic></source> <volume>15</volume>.<issue>1</issue> (<year>2014</year>), <comment>pp.</comment> <fpage>1929</fpage>&#x2013;<lpage>1958</lpage> (<comment>cit. on pp. 110, 127, 141, 176</comment>).</mixed-citation></ref>
<ref id="CIT374"><label>[Ste+14]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Heinz</given-names> <surname>Christian Steinhausen</surname></string-name>, <string-name><given-names>Dennis den</given-names> <surname>Brok</surname></string-name>, <string-name><given-names>Matthias B</given-names> <surname>Hullin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <article-title>&#x201C;Acquiring Bidirectional Texture Functions for Large-Scale Material Samples&#x201D;</article-title>. <comment>In</comment>: (<year>2014</year>) (<comment>cit. on pp. 30, 63</comment>).</mixed-citation></ref>
<ref id="CIT375"><label>[Ste+15a]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Heinz</given-names> <surname>Christian Steinhausen</surname></string-name>, <string-name><given-names>Dennis den</given-names> <surname>Brok</surname></string-name>, <string-name><given-names>Matthias B.</given-names> <surname>Hullin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <chapter-title>&#x201C;Extrapolating Large-Scale Material BTFs under Cross-Device Constraints&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Vision, Modeling &#x0026; Visualization</italic></source>. <role>Ed.</role> <comment>by</comment> <person-group person-group-type="editor"><string-name><given-names>David</given-names> <surname>Bommes</surname></string-name>, <string-name><given-names>Tobias</given-names> <surname>Ritschel</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="editor"><string-name><given-names>Thomas</given-names> <surname>Schultz</surname></string-name></person-group>. <publisher-name>The Eurographics Association</publisher-name>, <year>2015</year>, <comment>pp.</comment> <fpage>143</fpage>&#x2013;<lpage>150</lpage> (<comment>cit. on pp. 33, 104</comment>).</mixed-citation></ref>
<ref id="CIT376"><label>[Ste+15b]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Heinz Christian</given-names> <surname>Steinhausen</surname></string-name>, <string-name><given-names>Rodrigo</given-names> <surname>Mart&#x00ED;n</surname></string-name>, <string-name><given-names>Dennis den</given-names> <surname>Brok</surname></string-name>, <string-name><given-names>Matthias B.</given-names> <surname>Hullin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <chapter-title>&#x201C;Extrapolation of Bidirectional Texture Functions using Texture Synthesis guided by Photometric Normals&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Measuring, Modeling, and Reproducing Material Appearance II (SPIE 9398)</italic></source>. <volume>Vol. 9398</volume>. <issue>14</issue>. <publisher-name>San Francisco, USA</publisher-name>, <month>Feb</month>. <year>2015</year> (<comment>cit. on pp. 33, 104</comment>).</mixed-citation></ref>
<ref id="CIT377"><label>[Str+22]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Yannick</given-names> <surname>Str&#x00FC;mpler</surname></string-name>, <string-name><given-names>Janis</given-names> <surname>Postels</surname></string-name>, <string-name><given-names>Ren</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Luc</given-names> <surname>Van Gool</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Federico</given-names> <surname>Tombari</surname></string-name></person-group>. <chapter-title>&#x201C;Implicit Neural Representations for Image Compression&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2022</year>, <comment>pp.</comment> <fpage>74</fpage>&#x2013;<lpage>91</lpage> (<comment>cit. on p. 187</comment>).</mixed-citation></ref>
<ref id="CIT378"><label>[SGK21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Vadim</given-names> <surname>Sushko</surname></string-name>, <string-name><given-names>Juergen</given-names> <surname>Gall</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Anna</given-names> <surname>Khoreva</surname></string-name></person-group>. <article-title>&#x201C;One-Shot GAN: Learning to Generate Samples from Single Images and Videos&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2103.13389</italic></source> (<year>2021</year>) (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT379"><label>[SB08]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C&#x00E9;dric</given-names> <surname>Syllebranque</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Samuel</given-names> <surname>Boivin</surname></string-name></person-group>. <article-title>&#x201C;Estimation of Mechanical Parameters of Deformable Solids from Videos&#x201D;</article-title>. <comment>In</comment>: <source><italic>The Visual Computer</italic></source> <volume>24</volume>.<issue>11</issue> (<year>2008</year>), <comment>pp.</comment> <fpage>963</fpage>&#x2013;<lpage>972</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT380"><label>[Szt+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Alejandro</given-names> <surname>Sztrajman</surname></string-name>, <string-name><given-names>Gilles</given-names> <surname>Rainer</surname></string-name>, <string-name><given-names>Tobias</given-names> <surname>Ritschel</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Tim</given-names> <surname>Weyrich</surname></string-name></person-group>. <article-title>&#x201C;Neural BRDF Representation and Importance Sampling&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 40. 6</volume>. <publisher-name>Wiley Online Library</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>332</fpage>&#x2013;<lpage>346</lpage> (<comment>cit. on pp. 105, 111, 116, 186, 187, 193, 203, 210</comment>).</mixed-citation></ref>
<ref id="CIT381"><label>[Tak+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Towaki</given-names> <surname>Takikawa</surname></string-name>, <string-name><given-names>Joey</given-names> <surname>Litalien</surname></string-name>, <string-name><given-names>Kangxue</given-names> <surname>Yin</surname></string-name>, <string-name><given-names>Karsten</given-names> <surname>Kreis</surname></string-name>, <string-name><given-names>Charles</given-names> <surname>Loop</surname></string-name>, <string-name><given-names>Derek</given-names> <surname>Nowrouzezahrai</surname></string-name>, <string-name><given-names>Alec</given-names> <surname>Jacobson</surname></string-name>, <string-name><given-names>Morgan</given-names> <surname>McGuire</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Sanja</given-names> <surname>Fidler</surname></string-name></person-group>. <article-title>&#x201C;Neural Geometric Level of Detail: Real-time Rendering with Implicit 3D Shapes&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2021</year> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT382"><label>[Tam+11]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_259"/><person-group person-group-type="author"><string-name><given-names>Omer</given-names> <surname>Tamuz</surname></string-name>, <string-name><given-names>Ce</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Serge</given-names> <surname>Belongie</surname></string-name>, <string-name><given-names>Ohad</given-names> <surname>Shamir</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Adam</given-names> <surname>Tauman Kalai</surname></string-name></person-group>. <article-title>&#x201C;Adaptively Learning the Crowd Kernel&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <year>2011</year>, <comment>pp.</comment> <fpage>673</fpage>&#x2013;<lpage>680</lpage> (<comment>cit. on pp. 170, 171</comment>).</mixed-citation></ref>
<ref id="CIT383"><label>[Tan+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Matthew</given-names> <surname>Tancik</surname></string-name>, <string-name><given-names>Vincent</given-names> <surname>Casser</surname></string-name>, <string-name><given-names>Xinchen</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>Sabeek</given-names> <surname>Pradhan</surname></string-name>, <string-name><given-names>Ben</given-names> <surname>Mildenhall</surname></string-name>, <string-name><given-names>Pratul</given-names> <surname>Srinivasan</surname></string-name>, <string-name><given-names>Jonathan T.</given-names> <surname>Barron</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Henrik</given-names> <surname>Kretzschmar</surname></string-name></person-group>. <article-title>&#x201C;Block-NeRF: Scalable Large Scene Neural View Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv</italic></source>. <year>2022</year> (<comment>cit. on p. 187</comment>).</mixed-citation></ref>
<ref id="CIT384"><label>[Tan+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Matthew</given-names> <surname>Tancik</surname></string-name>, <string-name><given-names>Pratul P.</given-names> <surname>Srinivasan</surname></string-name>, <string-name><given-names>Ben</given-names> <surname>Mildenhall</surname></string-name>, <string-name><given-names>Sara</given-names> <surname>Fridovich-Keil</surname></string-name>, <string-name><given-names>Nithin</given-names> <surname>Raghavan</surname></string-name>, <string-name><given-names>Utkarsh</given-names> <surname>Singhal</surname></string-name>, <string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name>, <string-name><given-names>Jonathan T.</given-names> <surname>Barron</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ren</given-names> <surname>Ng</surname></string-name></person-group>. <article-title>&#x201C;Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> (<year>2020</year>) (<comment>cit. on pp. 75, 186, 193</comment>).</mixed-citation></ref>
<ref id="CIT385"><label>[Tew+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ayush</given-names> <surname>Tewari</surname></string-name>, <string-name><given-names>Justus</given-names> <surname>Thies</surname></string-name>, <string-name><given-names>Ben</given-names> <surname>Mildenhall</surname></string-name>, <string-name><given-names>Pratul</given-names> <surname>Srinivasan</surname></string-name>, <string-name><given-names>Edgar</given-names> <surname>Tretschk</surname></string-name>, <string-name><given-names>Wang</given-names> <surname>Yifan</surname></string-name>, <string-name><given-names>Christoph</given-names> <surname>Lassner</surname></string-name>, <string-name><given-names>Vincent</given-names> <surname>Sitzmann</surname></string-name>, <string-name><given-names>Ricardo</given-names> <surname>Martin-Brualla</surname></string-name>, <string-name><given-names>Stephen</given-names> <surname>Lombardi</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;Advances in Neural Rendering&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 41</volume>. <issue>2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2022</year>, <comment>pp.</comment> <fpage>703</fpage>&#x2013;<lpage>735</lpage> (<comment>cit. on pp. 187, 212</comment>).</mixed-citation></ref>
<ref id="CIT386"><label>[Tex+20a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ond&#x0159;ej</given-names> <surname>Texler</surname></string-name>, <string-name><given-names>David</given-names> <surname>Futschik</surname></string-name>, <string-name><given-names>Jakub</given-names> <surname>Fi&#x0161;er</surname></string-name>, <string-name><given-names>Michal</given-names> <surname>Luk&#x00E1;&#x010D;</surname></string-name>, <string-name><given-names>Jingwan</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Sy&#x0060;kora</surname></string-name></person-group>. <article-title>&#x201C;Arbitrary Style Transfer using Neurally-Guided Patch-Based Ssynthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computers &#x0026; Graphics</italic></source> <volume>87</volume> (<year>2020</year>), <comment>pp.</comment> <fpage>62</fpage>&#x2013;<lpage>71</lpage> (<comment>cit. on pp. 39, 45</comment>).</mixed-citation></ref>
<ref id="CIT387"><label>[Tex+20b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ond&#x0159;ej</given-names> <surname>Texler</surname></string-name>, <string-name><given-names>David</given-names> <surname>Futschik</surname></string-name>, <string-name><given-names>Michal</given-names> <surname>Ku&#x010D;era</surname></string-name>, <string-name><given-names>Ond&#x0159;ej</given-names> <surname>Jamri&#x0161;ka</surname></string-name>, <string-name><given-names>&#x0160;&#x00E1;rka</given-names> <surname>Sochorov&#x00E1;</surname></string-name>, <string-name><given-names>Menclei</given-names> <surname>Chai</surname></string-name>, <string-name><given-names>Sergey</given-names> <surname>Tulyakov</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Sy&#x0060;kora</surname></string-name></person-group>. <article-title>&#x201C;Interactive Video Stylization Using Few-Shot Patch-Based Training&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>39</volume>.<issue>4</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>73</fpage>&#x2013;<lpage>1</lpage> (<comment>cit. on pp. 30, 31, 33, 37, 39, 46, 53, 58, 61, 107, 108, 128, 129</comment>).</mixed-citation></ref>
<ref id="CIT388"><label>[TOB15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Lucas</given-names> <surname>Theis</surname></string-name>, <string-name><given-names>A&#x00E4;ron</given-names> <surname>van den Oord</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Matthias</given-names> <surname>Bethge</surname></string-name></person-group>. <article-title>&#x201C;A Note on the Evaluation of Generative Models&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1511.01844</italic></source> (<year>2015</year>) (<comment>cit. on p. 41</comment>).</mixed-citation></ref>
<ref id="CIT389"><label>[Tom94]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Shoji</given-names> <surname>Tominaga</surname></string-name></person-group>. <article-title>&#x201C;Dichromatic reflection models for a variety of materials&#x201D;</article-title>. <comment>In</comment>: <source><italic>Color Research &#x0026; Application</italic></source> <volume>19</volume>.<issue>4</issue> (<year>1994</year>), <comment>pp.</comment> <fpage>277</fpage>&#x2013;<lpage>285</lpage> (<comment>cit. on p. 10</comment>).</mixed-citation></ref>
<ref id="CIT390"><label>[Ton+02]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xin</given-names> <surname>Tong</surname></string-name>, <string-name><given-names>Jingdan</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Ligang</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Xi</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Baining</given-names> <surname>Guo</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Heung-Yeung</given-names> <surname>Shum</surname></string-name></person-group>. <article-title>&#x201C;Synthesis of Bidirectional Texture Functions on Arbitrary Surfaces&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>21</volume>.<issue>3</issue> (<year>2002</year>), <comment>pp.</comment> <fpage>665</fpage>&#x2013;<lpage>672</lpage> (<comment>cit. on p. 103</comment>).</mixed-citation></ref>
<ref id="CIT391"><label>[TS67]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_260"/><person-group person-group-type="author"><string-name><given-names>Kenneth E</given-names> <surname>Torrance</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ephraim M</given-names> <surname>Sparrow</surname></string-name></person-group>. <article-title>&#x201C;Theory for Off-Specular Reflection from Roughened Surfaces&#x201D;</article-title>. <comment>In</comment>: <source><italic>Josa</italic></source> <volume>57</volume>.<issue>9</issue> (<year>1967</year>), <comment>pp.</comment> <fpage>1105</fpage>&#x2013;<lpage>1114</lpage> (<comment>cit. on p. 11</comment>).</mixed-citation></ref>
<ref id="CIT392"><label>[Tu+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Peihan</given-names> <surname>Tu</surname></string-name>, <string-name><given-names>Li-Yi</given-names> <surname>Wei</surname></string-name>, <string-name><given-names>Koji</given-names> <surname>Yatani</surname></string-name>, <string-name><given-names>Takeo</given-names> <surname>Igarashi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Matthias</given-names> <surname>Zwicker</surname></string-name></person-group>. <article-title>&#x201C;Continuous Curve Textures&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>39</volume>.<issue>6</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>16</lpage> (<comment>cit. on p. 72</comment>).</mixed-citation></ref>
<ref id="CIT393"><label>[TK74]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Amos</given-names> <surname>Tversky</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Daniel</given-names> <surname>Kahneman</surname></string-name></person-group>. <article-title>&#x201C;Judgment under Uncertainty: Heuristics and Biases: Biases in Judgments Reveal Some Heuristics of Thinking Under Uncertainty.&#x201D;</article-title> <comment>In</comment>: <source><italic>Science</italic></source> <volume>185</volume>.<issue>4157</issue> (<year>1974</year>), <comment>pp.</comment> <fpage>1124</fpage>&#x2013;<lpage>1131</lpage> (<comment>cit. on p. 169</comment>).</mixed-citation></ref>
<ref id="CIT394"><label>[Uly+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dmitry</given-names> <surname>Ulyanov</surname></string-name>, <string-name><given-names>Vadim</given-names> <surname>Lebedev</surname></string-name>, <string-name><given-names>Andrea</given-names> <surname>Vedaldi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Victor S</given-names> <surname>Lempitsky</surname></string-name></person-group>. <article-title>&#x201C;Texture Networks: Feed-Forward Synthesis of Textures and Stylized Images.&#x201D;</article-title> <comment>In</comment>: <source><italic>International Conference of Machine Learning (ICML)</italic></source>. <volume>Vol. 1</volume>. <issue>2</issue>. <year>2016</year>, <comment>p.</comment> <fpage>4</fpage> (<comment>cit. on p. 76</comment>).</mixed-citation></ref>
<ref id="CIT395"><label>[UVL18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dmitry</given-names> <surname>Ulyanov</surname></string-name>, <string-name><given-names>Andrea</given-names> <surname>Vedaldi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Victor</given-names> <surname>Lempitsky</surname></string-name></person-group>. <article-title>&#x201C;Deep Image Prior&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <month>June</month> <year>2018</year> (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT396"><label>[UVL16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dmitry</given-names> <surname>Ulyanov</surname></string-name>, <string-name><given-names>Andrea</given-names> <surname>Vedaldi</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Victor</given-names> <surname>Lempitsky</surname></string-name></person-group>. <article-title>&#x201C;Instance Normalization: The Missing Ingredient for Fast Stylization&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:1607.08022</italic></source> (<year>2016</year>) (<comment>cit. on pp. 77, 94</comment>).</mixed-citation></ref>
<ref id="CIT397"><label>[Van11]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dietger G</given-names> <surname>Van Antwerpen</surname></string-name></person-group>. <article-title>&#x201C;Unbiased Physically Nased Rendering on the GPU&#x201D;</article-title>. <source>MA thesis. Electrical Engineering, Mathematics and Computer Science</source>, <year>2011</year> (<comment>cit. on p. 195</comment>).</mixed-citation></ref>
<ref id="CIT398"><label>[Vas+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ashish</given-names> <surname>Vaswani</surname></string-name>, <string-name><given-names>Noam</given-names> <surname>Shazeer</surname></string-name>, <string-name><given-names>Niki</given-names> <surname>Parmar</surname></string-name>, <string-name><given-names>Jakob</given-names> <surname>Uszkoreit</surname></string-name>, <string-name><given-names>Llion</given-names> <surname>Jones</surname></string-name>, <string-name><given-names>Aidan N</given-names> <surname>Gomez</surname></string-name>, <string-name><given-names>&#x0141;ukasz</given-names> <surname>Kaiser</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Illia</given-names> <surname>Polosukhin</surname></string-name></person-group>. <article-title>&#x201C;Attention Is All You Need&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source> <volume>30</volume> (<year>2017</year>) (<comment>cit. on p. 164</comment>).</mixed-citation></ref>
<ref id="CIT399"><label>[VG95]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Eric</given-names> <surname>Veach</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Leonidas J.</given-names> <surname>Guibas</surname></string-name></person-group>. <chapter-title>&#x201C;Optimally Combining Sampling Techniques for Monte Carlo Rendering&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques</italic></source>. <comment>SIGGRAPH &#x2019;95</comment>. <publisher-name>Association for Computing Machinery</publisher-name>, <year>1995</year> (<comment>cit. on p. 184</comment>).</mixed-citation></ref>
<ref id="CIT400"><label>[VPS21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Giuseppe</given-names> <surname>Vecchio</surname></string-name>, <string-name><given-names>Simone</given-names> <surname>Palazzo</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Concetto</given-names> <surname>Spampinato</surname></string-name></person-group>. <article-title>&#x201C;SurfaceNet: Adversarial SVBRDF Estimation from a Single Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>12840</fpage>&#x2013;<lpage>12848</lpage> (<comment>cit. on pp. 17, 76, 123, 125, 128, 129, 132, 258</comment>).</mixed-citation></ref>
<ref id="CIT401"><label>[Ver+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Dor</given-names> <surname>Verbin</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Hedman</surname></string-name>, <string-name><given-names>Ben</given-names> <surname>Mildenhall</surname></string-name>, <string-name><given-names>Todd</given-names> <surname>Zickler</surname></string-name>, <string-name><given-names>Jonathan T.</given-names> <surname>Barron</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Pratul P.</given-names> <surname>Srinivasan</surname></string-name></person-group>. <article-title>&#x201C;Ref-NeRF: <target target-type="page" id="pges_261"/>Structured View-Dependent Appearance for Neural Radiance Fields&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2022</year> (<comment>cit. on pp. 187, 212</comment>).</mixed-citation></ref>
<ref id="CIT402"><label>[VZH20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yael</given-names> <surname>Vinker</surname></string-name>, <string-name><given-names>Nir</given-names> <surname>Zabari</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yedid</given-names> <surname>Hoshen</surname></string-name></person-group>. <article-title>&#x201C;Training End-to-end Single Image Generators without GANs&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2004.06014</italic></source> (<year>2020</year>) (<comment>cit. on p. 75</comment>).</mixed-citation></ref>
<ref id="CIT403"><label>[VMF09]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Pascal</given-names> <surname>Volino</surname></string-name>, <string-name><given-names>Nadia</given-names> <surname>Magnenat-Thalmann</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Francois</given-names> <surname>Faure</surname></string-name></person-group>. <article-title>&#x201C;A Simple Approach to Nonlinear Tensile Stiffness for Accurate Cloth Simulation&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>28</volume>.<issue>4</issue> (<year>2009</year>), <comment>Article&#x2013;No</comment> (<comment>cit. on pp. 158, 161, 162</comment>).</mixed-citation></ref>
<ref id="CIT404"><label>[Wal+14]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ingo</given-names> <surname>Wald</surname></string-name>, <string-name><given-names>Sven</given-names> <surname>Woop</surname></string-name>, <string-name><given-names>Carsten</given-names> <surname>Benthin</surname></string-name>, <string-name><given-names>Gregory S.</given-names> <surname>Johnson</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Manfred</given-names> <surname>Ernst</surname></string-name></person-group>. <article-title>&#x201C;Embree: A Kernel Framework for Efficient CPU Ray Tracing&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>33</volume>.<issue>4</issue> (<year>2014</year>) (<comment>cit. on p. 195</comment>).</mixed-citation></ref>
<ref id="CIT405"><label>[Wal+07]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Bruce</given-names> <surname>Walter</surname></string-name>, <string-name><given-names>Stephen R</given-names> <surname>Marschner</surname></string-name>, <string-name><given-names>Hongsong</given-names> <surname>Li</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Kenneth E</given-names> <surname>Torrance</surname></string-name></person-group>. <article-title>&#x201C;Microfacet Models for Refraction through Rough Surfaces&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the 18th Eurographics Conference on Rendering Techniques</italic></source>. <year>2007</year>, <comment>pp.</comment> <fpage>195</fpage>&#x2013;<lpage>206</lpage> (<comment>cit. on p. 127</comment>).</mixed-citation></ref>
<ref id="CIT406"><label>[WZY22]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Cairong</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Yiming</given-names> <surname>Zhu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Chun</given-names> <surname>Yuan</surname></string-name></person-group>. <chapter-title>&#x201C;Diverse Image Inpainting with Normalizing Flow&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2022</year>, <comment>pp.</comment> <fpage>53</fpage>&#x2013;<lpage>69</lpage> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT407"><label>[Wan+22a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Can</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Menglei</given-names> <surname>Chai</surname></string-name>, <string-name><given-names>Mingming</given-names> <surname>He</surname></string-name>, <string-name><given-names>Dongdong</given-names> <surname>Chen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name></person-group>. <article-title>&#x201C;Clip-nerf: Text-and-image Driven Manipulation of Neural Radiance Fields&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>3835</fpage>&#x2013;<lpage>3844</lpage> (<comment>cit. on p. 116</comment>).</mixed-citation></ref>
<ref id="CIT408"><label>[Wan+22b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Chen</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Xiang</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Jiawei</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Liang</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Xiao</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>Xin</given-names> <surname>Ning</surname></string-name>, <string-name><given-names>Jun</given-names> <surname>Zhou</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Edwin</given-names> <surname>Hancock</surname></string-name></person-group>. <article-title>&#x201C;Uncertainty Estimation for Stereo Matching Based on Evidential Deep Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>Pattern Recognition</italic></source> <volume>124</volume> (<year>2022</year>), <comment>p.</comment> <fpage>108498</fpage> (<comment>cit. on p. 130</comment>).</mixed-citation></ref>
<ref id="CIT409"><label>[Wan+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Guotai</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Wenqi</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Maria A</given-names> <surname>Zuluaga</surname></string-name>, <string-name><given-names>Rosalind</given-names> <surname>Pratt</surname></string-name>, <string-name><given-names>Premal A</given-names> <surname>Patel</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Aertsen</surname></string-name>, <string-name><given-names>Tom</given-names> <surname>Doel</surname></string-name>, <string-name><given-names>Anna L</given-names> <surname>David</surname></string-name>, <string-name><given-names>Jan</given-names> <surname>Deprest</surname></string-name>, <string-name><given-names>S&#x00E9;bastien</given-names> <surname>Ourselin</surname></string-name></person-group>, <etal>et al.</etal> <article-title>&#x201C;Interactive Medical Image Segmentation using Deep Learning with Image-specific Fine Tuning&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE transactions on medical imaging</italic></source> <volume>37</volume>.<issue>7</issue> (<year>2018</year>), <comment>pp.</comment> <fpage>1562</fpage>&#x2013;<lpage>1573</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT410"><label>[WOR11]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Huamin</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>James F</given-names> <surname>O&#x2019;Brien</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ravi</given-names> <surname>Ramamoorthi</surname></string-name></person-group>. <article-title>&#x201C;Data-driven Elastic Models for Cloth: Modeling and Measurement&#x201D;</article-title>. <target target-type="page" id="pges_262"/><comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>30</volume>.<issue>4</issue> (<year>2011</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>12</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT411"><label>[Wan+09]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jiaping</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Peiran</given-names> <surname>Ren</surname></string-name>, <string-name><given-names>Minmin</given-names> <surname>Gong</surname></string-name>, <string-name><given-names>John</given-names> <surname>Snyder</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Baining</given-names> <surname>Guo</surname></string-name></person-group>. <article-title>&#x201C;All-Frequency Rendering of Dynamic, Spatially-Varying Reflectance&#x201D;</article-title>. <comment>In</comment>: <volume>28</volume>.<issue>5</issue> (<month>Dec</month>. <year>2009</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>10</lpage> (<comment>cit. on p. 13</comment>).</mixed-citation></ref>
<ref id="CIT412"><label>[Wan+22c]</label> <mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Jiayi</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Diogo</given-names> <surname>Luvizon</surname></string-name>, <string-name><given-names>Franziska</given-names> <surname>Mueller</surname></string-name>, <string-name><given-names>Florian</given-names> <surname>Bernard</surname></string-name>, <string-name><given-names>Adam</given-names> <surname>Kortylewski</surname></string-name>, <string-name><given-names>Dan</given-names> <surname>Casas</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Christian</given-names> <surname>Theobalt</surname></string-name></person-group>. <article-title>&#x201C;HandFlow: Quantifying View-Dependent 3D Ambiguity in Two-Hand Reconstruction with Normalizing Flow&#x201D;</article-title>. <comment>In</comment>: <conf-name>International Symposium on Vision, Modeling, and Visualization</conf-name>. <year>2022</year> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT413"><label>[Wan+20a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jiayun</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Yubei</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Rudrasis</given-names> <surname>Chakraborty</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Stella X.</given-names> <surname>Yu</surname></string-name></person-group>. <article-title>&#x201C;Orthogonal Convolutional Neural Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <month>June</month> <year>2020</year> (<comment>cit. on p. 93</comment>).</mixed-citation></ref>
<ref id="CIT414"><label>[Wan+22d]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Liao</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Jiakai</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Xinhang</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Fuqiang</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Yanshun</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Yingliang</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Minye</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Jingyi</given-names> <surname>Yu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Lan</given-names> <surname>Xu</surname></string-name></person-group>. <article-title>&#x201C;Fourier Plenoctrees for Dynamic Radiance Field Rendering in Real-time&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2022</year>, <comment>pp.</comment> <fpage>13524</fpage>&#x2013;<lpage>13534</lpage> (<comment>cit. on pp. 203, 212</comment>).</mixed-citation></ref>
<ref id="CIT415"><label>[Wan+20b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sinong</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Belinda Z</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Madian</given-names> <surname>Khabsa</surname></string-name>, <string-name><given-names>Han</given-names> <surname>Fang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hao</given-names> <surname>Ma</surname></string-name></person-group>. <article-title>&#x201C;Linformer: Self-attention with Linear Complexity&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2006.04768</italic></source> (<year>2020</year>) (<comment>cit. on pp. 124, 127, 141, 144</comment>).</mixed-citation></ref>
<ref id="CIT416"><label>[Wan+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ting-Chun</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Ming-Yu</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Andrew</given-names> <surname>Tao</surname></string-name>, <string-name><given-names>Guilin</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Bryan</given-names> <surname>Catanzaro</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Kautz</surname></string-name></person-group>. <article-title>&#x201C;Few-shot Video-to-Video Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>5013</fpage>&#x2013;<lpage>5024</lpage> (<comment>cit. on p. 37</comment>).</mixed-citation></ref>
<ref id="CIT417"><label>[Wan+20c]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xiaofei</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Yiwen</given-names> <surname>Han</surname></string-name>, <string-name><given-names>Victor CM</given-names> <surname>Leung</surname></string-name>, <string-name><given-names>Dusit</given-names> <surname>Niyato</surname></string-name>, <string-name><given-names>Xueqiang</given-names> <surname>Yan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Xu</given-names> <surname>Chen</surname></string-name></person-group>. <article-title>&#x201C;Convergence of Edge Computing and Deep Learning: A Comprehensive Survey&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Communications Surveys &#x0026; Tutorials</italic></source> <volume>22</volume>.<issue>2</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>869</fpage>&#x2013;<lpage>904</lpage> (<comment>cit. on p. 212</comment>).</mixed-citation></ref>
<ref id="CIT418"><label>[Wan+20d]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yaqing</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Quanming</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>James T</given-names> <surname>Kwok</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Lionel M</given-names> <surname>Ni</surname></string-name></person-group>. <article-title>&#x201C;Generalizing From a Few Examples: A Survey on Few-Shot Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Computing Surveys (CSUR)</italic></source> <volume>53</volume>.<issue>3</issue> (<year>2020</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>34</lpage> (<comment>cit. on p. 37</comment>).</mixed-citation></ref>
<ref id="CIT419"><label>[Wan+04]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhou</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Alan C</given-names> <surname>Bovik</surname></string-name>, <string-name><given-names>Hamid R</given-names> <surname>Sheikh</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Eero P</given-names> <surname>Simoncelli</surname></string-name></person-group>. <article-title>&#x201C;Image Quality Assessment: From Error Visibility to Structural Similarity&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Image <target target-type="page" id="pges_263"/>Processing</italic></source> <volume>13</volume>.<issue>4</issue> (<year>2004</year>), <comment>pp.</comment> <fpage>600</fpage>&#x2013;<lpage>612</lpage> (<comment>cit. on pp. 50, 54, 85, 86, 90, 92, 160, 198, 199, 202</comment>).</mixed-citation></ref>
<ref id="CIT420"><label>[Wan+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zian</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Jonah</given-names> <surname>Philion</surname></string-name>, <string-name><given-names>Sanja</given-names> <surname>Fidler</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Jan</given-names> <surname>Kautz</surname></string-name></person-group>. <article-title>&#x201C;Learning Indoor Inverse Rendering with 3D Spatially-Varying Lighting&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>12538</fpage>&#x2013;<lpage>12547</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT421"><label>[Wan+23]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zian</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Tianchang</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>Jun</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>Shengyu</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Jacob</given-names> <surname>Munkberg</surname></string-name>, <string-name><given-names>Jon</given-names> <surname>Hasselgren</surname></string-name>, <string-name><given-names>Zan</given-names> <surname>Gojcic</surname></string-name>, <string-name><given-names>Wenzheng</given-names> <surname>Chen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Sanja</given-names> <surname>Fidler</surname></string-name></person-group>. <article-title>&#x201C;Neural Fields meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2023</year> (<comment>cit. on pp. 203, 212</comment>).</mixed-citation></ref>
<ref id="CIT422"><label>[WGK14]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Michael</given-names> <surname>Weinmann</surname></string-name>, <string-name><given-names>Juergen</given-names> <surname>Gall</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Reinhard</given-names> <surname>Klein</surname></string-name></person-group>. <chapter-title>&#x201C;Material Classification Based on Training Data Synthesized Using a BTF Database&#x201D;</chapter-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer International Publishing</publisher-name>, <year>2014</year>, <fpage>pp.</fpage> <fpage>156</fpage>&#x2013;<lpage>171</lpage> (<comment>cit. on pp. 50, 103, 108, 111, 151</comment>).</mixed-citation></ref>
<ref id="CIT423"><label>[WKW16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Karl</given-names> <surname>Weiss</surname></string-name>, <string-name><given-names>Taghi M</given-names> <surname>Khoshgoftaar</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>DingDing</given-names> <surname>Wang</surname></string-name></person-group>. <article-title>&#x201C;A Survey of Transfer Learning&#x201D;</article-title>. <comment>In</comment>: <source><italic>Journal of Big data</italic></source> <volume>3</volume>.<issue>1</issue> (<year>2016</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>40</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT424"><label>[Wen+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tao</given-names> <surname>Wen</surname></string-name>, <string-name><given-names>Beibei</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Lei</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Jie</given-names> <surname>Guo</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Nicolas</given-names> <surname>Holzschuch</surname></string-name></person-group>. <article-title>&#x201C;SVBRDF Recovery from a Single Image with Highlights Using a Pre-trained Generative Adversarial Network&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <publisher-name>Wiley Online Library</publisher-name>. <year>2022</year> (<comment>cit. on p. 125</comment>).</mixed-citation></ref>
<ref id="CIT425"><label>[Woo+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sanghyun</given-names> <surname>Woo</surname></string-name>, <string-name><given-names>Jongchan</given-names> <surname>Park</surname></string-name>, <string-name><given-names>Joon-Young</given-names> <surname>Lee</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>In So</given-names> <surname>Kweon</surname></string-name></person-group>. <article-title>&#x201C;Cbam: Convolutional block attention module&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <year>2018</year>, <comment>pp.</comment> <fpage>3</fpage>&#x2013;<lpage>19</lpage> (<comment>cit. on pp. 109, 119, 128, 141, 145, 164</comment>).</mixed-citation></ref>
<ref id="CIT426"><label>[WTX22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Qiling</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Jianchao</given-names> <surname>Tan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Kun</given-names> <surname>Xu</surname></string-name></person-group>. <article-title>&#x201C;PaletteNeRF: Palette-based Color Editing for NeRFs&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2212.12871</italic></source> (<year>2022</year>) (<comment>cit. on p. 116</comment>).</mixed-citation></ref>
<ref id="CIT427"><label>[WH18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yuxin</given-names> <surname>Wu</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Kaiming</given-names> <surname>He</surname></string-name></person-group>. <article-title>&#x201C;Group Normalization&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <year>2018</year>, <comment>pp.</comment> <fpage>3</fpage>&#x2013;<lpage>19</lpage> (<comment>cit. on pp. 64, 67, 69, 127, 141, 144</comment>).</mixed-citation></ref>
<ref id="CIT428"><label>[Xie+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yiheng</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>Towaki</given-names> <surname>Takikawa</surname></string-name>, <string-name><given-names>Shunsuke</given-names> <surname>Saito</surname></string-name>, <string-name><given-names>Or</given-names> <surname>Litany</surname></string-name>, <string-name><given-names>Shiqin</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>Numair</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>Federico</given-names> <surname>Tombari</surname></string-name>, <string-name><given-names>James</given-names> <surname>Tompkin</surname></string-name>, <string-name><given-names>Vincent</given-names> <surname>Sitzmann</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Srinath</given-names> <surname>Sridhar</surname></string-name></person-group>. <article-title>&#x201C;Neural Fields in Visual Computing and Beyond&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source> (<year>2022</year>) (<comment>cit. on pp. 187, 212</comment>).</mixed-citation></ref>
<ref id="CIT429"><label>[Xu+13]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_264"/><person-group person-group-type="author"><string-name><given-names>Kun</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Wei-Lun</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>Zhao</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Dan-Yong</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Run-Dong</given-names> <surname>Wu</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Shi-Min</given-names> <surname>Hu</surname></string-name></person-group>. <article-title>&#x201C;Anisotropic Spherical Gaussians&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>32</volume>.<issue>6</issue> (<year>2013</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>11</lpage> (<comment>cit. on pp. 185, 196, 199, 202</comment>).</mixed-citation></ref>
<ref id="CIT430"><label>[Yan+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Guandao</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Serge</given-names> <surname>Belongie</surname></string-name>, <string-name><given-names>Bharath</given-names> <surname>Hariharan</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Vladlen</given-names> <surname>Koltun</surname></string-name></person-group>. <article-title>&#x201C;Geometry processing with Neural Fields&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <volume>Vol. 34</volume>. <year>2021</year>, <comment>pp.</comment> <fpage>22483</fpage>&#x2013;<lpage>22497</lpage> (<comment>cit. on p. 186</comment>).</mixed-citation></ref>
<ref id="CIT431"><label>[YLL17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Shan</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Junbang</given-names> <surname>Liang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ming C</given-names> <surname>Lin</surname></string-name></person-group>. <article-title>&#x201C;Learning-based Cloth Material Recovery from Video&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE International Conference on Computer Vision (ICCV)</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>4383</fpage>&#x2013;<lpage>4393</lpage> (<comment>cit. on pp. 156, 159</comment>).</mixed-citation></ref>
<ref id="CIT432"><label>[YL15]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Shan</given-names> <surname>Yang</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ming C</given-names> <surname>Lin</surname></string-name></person-group>. <article-title>&#x201C;Materialcloning: Acquiring Elasticity Parameters from Images for Medical Applications&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Visualization and Computer Graphics</italic></source> <volume>22</volume>.<issue>9</issue> (<year>2015</year>), <comment>pp.</comment> <fpage>2122</fpage>&#x2013;<lpage>2135</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT433"><label>[Yan+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Shan</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Zherong</given-names> <surname>Pan</surname></string-name>, <string-name><given-names>Tanya</given-names> <surname>Amert</surname></string-name>, <string-name><given-names>Ke</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Licheng</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>Tamara</given-names> <surname>Berg</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ming C</given-names> <surname>Lin</surname></string-name></person-group>. <article-title>&#x201C;Physics-inspired Garment Recovery from a Single-view Image&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>37</volume>.<issue>5</issue> (<year>2018</year>), <comment>pp.</comment> <fpage>1</fpage>&#x2013;<lpage>14</lpage> (<comment>cit. on p. 158</comment>).</mixed-citation></ref>
<ref id="CIT434"><label>[Yao+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yao</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>Jingyang</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Jingbo</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Yihang</given-names> <surname>Qu</surname></string-name>, <string-name><given-names>Tian</given-names> <surname>Fang</surname></string-name>, <string-name><given-names>David</given-names> <surname>McKinnon</surname></string-name>, <string-name><given-names>Yanghai</given-names> <surname>Tsin</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Long</given-names> <surname>Quan</surname></string-name></person-group>. <article-title>&#x201C;NeILF: Neural Incident Light Field for Material and Lighting Estimation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <year>2022</year> (<comment>cit. on p. 187</comment>).</mixed-citation></ref>
<ref id="CIT435"><label>[Ye+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Weicai</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>Shuo</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Chong</given-names> <surname>Bao</surname></string-name>, <string-name><given-names>Hujun</given-names> <surname>Bao</surname></string-name>, <string-name><given-names>Marc</given-names> <surname>Pollefeys</surname></string-name>, <string-name><given-names>Zhaopeng</given-names> <surname>Cui</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Guofeng</given-names> <surname>Zhang</surname></string-name></person-group>. <article-title>&#x201C;Intrinsicnerf: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis&#x201D;</article-title>. <comment>In</comment>: <source><italic>arXiv preprint arXiv:2210.00647</italic></source> (<year>2022</year>) (<comment>cit. on p. 116</comment>).</mixed-citation></ref>
<ref id="CIT436"><label>[Ye+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wenjie</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>Yue</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Pieter</given-names> <surname>Peers</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Baining</given-names> <surname>Guo</surname></string-name></person-group>. <article-title>&#x201C;Deep Reflectance Scanning: Recovering Spatially-varying Material Appearance from a Flash-lit Video Sequence&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 40</volume>. <issue>6</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>409</fpage>&#x2013;<lpage>427</lpage> (<comment>cit. on p. 125</comment>).</mixed-citation></ref>
<ref id="CIT437"><label>[Ye+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wenjie</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>Xiao</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Yue</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>Pieter</given-names> <surname>Peers</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Xin</given-names> <surname>Tong</surname></string-name></person-group>. <article-title>&#x201C;Single Image Surface Appearance Modeling with Self-Augmented CNNs and Inexact Supervision&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 37</volume>. <issue>7</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2018</year>, <comment>pp.</comment> <fpage>201</fpage>&#x2013;<lpage>211</lpage> (<comment>cit. on pp. 34, 125</comment>).</mixed-citation></ref>
<ref id="CIT438"><label>[Yep20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Tyler</given-names> <surname>Yep</surname></string-name></person-group>. <source><italic>torchinfo</italic></source>. <month>Mar</month>. <year>2020</year> (<comment>cit. on p. 113</comment>).</mixed-citation></ref>
<ref id="CIT439"><label>[YS19]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_265"/><person-group person-group-type="author"><string-name><given-names>Ye</given-names> <surname>Yu</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>William AP</given-names> <surname>Smith</surname></string-name></person-group>. <article-title>&#x201C;InverseRenderNet: Learning Single Image Inverse Rendering&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>3155</fpage>&#x2013;<lpage>3164</lpage> (<comment>cit. on pp. 78, 203</comment>).</mixed-citation></ref>
<ref id="CIT440"><label>[YS21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Ye</given-names> <surname>Yu</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>William AP</given-names> <surname>Smith</surname></string-name></person-group>. <article-title>&#x201C;Outdoor Inverse Rendering from a Single Image using Multiview Self-supervision&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Pattern Analysis and Machine Intelligence</italic></source> <volume>44</volume>.<issue>7</issue> (<year>2021</year>), <comment>pp.</comment> <fpage>3659</fpage>&#x2013;<lpage>3675</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT441"><label>[Yun+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Sangdoo</given-names> <surname>Yun</surname></string-name>, <string-name><given-names>Dongyoon</given-names> <surname>Han</surname></string-name>, <string-name><given-names>Seong Joon</given-names> <surname>Oh</surname></string-name>, <string-name><given-names>Sanghyuk</given-names> <surname>Chun</surname></string-name>, <string-name><given-names>Junsuk</given-names> <surname>Choe</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Youngjoon</given-names> <surname>Yoo</surname></string-name></person-group>. <article-title>&#x201C;Cutmix: Regularization strategy to train strong classifiers with localizable features&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>6023</fpage>&#x2013;<lpage>6032</lpage> (<comment>cit. on p. 93</comment>).</mixed-citation></ref>
<ref id="CIT442"><label>[Zha+19a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Bo</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Mingming</given-names> <surname>He</surname></string-name>, <string-name><given-names>Jing</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>Pedro V</given-names> <surname>Sander</surname></string-name>, <string-name><given-names>Lu</given-names> <surname>Yuan</surname></string-name>, <string-name><given-names>Amine</given-names> <surname>Bermak</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Dong</given-names> <surname>Chen</surname></string-name></person-group>. <article-title>&#x201C;Deep Exemplar-based Video Colorization&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>8052</fpage>&#x2013;<lpage>8061</lpage> (<comment>cit. on p. 33</comment>).</mixed-citation></ref>
<ref id="CIT443"><label>[Zha+19b]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Han</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Ian</given-names> <surname>Goodfellow</surname></string-name>, <string-name><given-names>Dimitris</given-names> <surname>Metaxas</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Augustus</given-names> <surname>Odena</surname></string-name></person-group>. <article-title>&#x201C;Self-attention Generative Adversarial Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Machine Learning (ICML)</italic></source>. <publisher-name>PMLR</publisher-name>. <year>2019</year>, <comment>pp.</comment> <fpage>7354</fpage>&#x2013;<lpage>7363</lpage> (<comment>cit. on pp. 164, 166, 176</comment>).</mixed-citation></ref>
<ref id="CIT444"><label>[ZDN16]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Hang</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Kristin</given-names> <surname>Dana</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ko</given-names> <surname>Nishino</surname></string-name></person-group>. <article-title>&#x201C;Friction from Reflectance: Deep Reflectance Codes for Predicting Physical Surface Properties from One-shot In-field Reflectance&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the European Conference on Computer Vision (ECCV)</italic></source>. <publisher-name>Springer</publisher-name>. <year>2016</year>, <comment>pp.</comment> <fpage>808</fpage>&#x2013;<lpage>824</lpage> (<comment>cit. on p. 159</comment>).</mixed-citation></ref>
<ref id="CIT445"><label>[Zha+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jingzhao</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Tianxing</given-names> <surname>He</surname></string-name>, <string-name><given-names>Suvrit</given-names> <surname>Sra</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Ali</given-names> <surname>Jadbabaie</surname></string-name></person-group>. <article-title>&#x201C;Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity&#x201D;</article-title>. <comment>In</comment>: <source><italic>International Conference on Learning Representations (ICLR)</italic></source>. <year>2020</year> (<comment>cit. on pp. 110, 194</comment>).</mixed-citation></ref>
<ref id="CIT446"><label>[Zha+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Kai</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Fujun</given-names> <surname>Luan</surname></string-name>, <string-name><given-names>Qianqian</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Kavita</given-names> <surname>Bala</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Noah</given-names> <surname>Snavely</surname></string-name></person-group>. <article-title>&#x201C;Physg: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2021</year>, <comment>pp.</comment> <fpage>5453</fpage>&#x2013;<lpage>5462</lpage> (<comment>cit. on p. 203</comment>).</mixed-citation></ref>
<ref id="CIT447"><label>[ZL12]</label> <mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>Lin</given-names> <surname>Zhang</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hongyu</given-names> <surname>Li</surname></string-name></person-group>. <article-title>&#x201C;SR-SIM: A Fast and High Performance IQA Index Based on Spectral Residual&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE International Conference on Image Processing</italic></source>. <publisher-name>IEEE</publisher-name>. <year>2012</year>, <comment>pp.</comment> <fpage>1473</fpage>&#x2013;<lpage>1476</lpage> (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT448"><label>[ZSL14]</label> <mixed-citation publication-type="journal"><target target-type="page" id="pges_266"/><person-group person-group-type="author"><string-name><given-names>Lin</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Ying</given-names> <surname>Shen</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hongyu</given-names> <surname>Li</surname></string-name></person-group>. <article-title>&#x201C;VSI: A Visual Saliency-induced Index for Perceptual Image Quality Assessment&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Image processing</italic></source> <volume>23</volume>.<issue>10</issue> (<year>2014</year>), <comment>pp.</comment> <fpage>4270</fpage>&#x2013;<lpage>4281</lpage> (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT449"><label>[Zha+11]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Lin</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Lei</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Xuanqin</given-names> <surname>Mou</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>David</given-names> <surname>Zhang</surname></string-name></person-group>. <article-title>&#x201C;FSIM: A Feature Similarity Index for Image Quality Assessment&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Image Processing</italic></source> <volume>20</volume>.<issue>8</issue> (<year>2011</year>), <comment>pp.</comment> <fpage>2378</fpage>&#x2013;<lpage>2386</lpage> (<comment>cit. on p. 160</comment>).</mixed-citation></ref>
<ref id="CIT450"><label>[Zha+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Richard</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Phillip</given-names> <surname>Isola</surname></string-name>, <string-name><given-names>Alexei A.</given-names> <surname>Efros</surname></string-name>, <string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Oliver</given-names> <surname>Wang</surname></string-name></person-group>. <article-title>&#x201C;The Unreasonable Effectiveness of Deep Features as a Perceptual Metric&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2018</year>, <comment>pp.</comment> <fpage>586</fpage>&#x2013;<lpage>595</lpage> (<comment>cit. on pp. 32, 39, 41, 49, 50, 54, 57, 67, 85, 86, 90, 110, 128, 142, 157, 160, 169, 170, 178, 198, 199, 202</comment>).</mixed-citation></ref>
<ref id="CIT451"><label>[ZJK20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Hengshuang</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Jiaya</given-names> <surname>Jia</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Vladlen</given-names> <surname>Koltun</surname></string-name></person-group>. <article-title>&#x201C;Exploring Self-attention for Image Recognition&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2020</year>, <comment>pp.</comment> <fpage>10076</fpage>&#x2013;<lpage>10085</lpage> (<comment>cit. on p. 164</comment>).</mixed-citation></ref>
<ref id="CIT452"><label>[Zho+20]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Zhun</given-names> <surname>Zhong</surname></string-name>, <string-name><given-names>Liang</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>Guoliang</given-names> <surname>Kang</surname></string-name>, <string-name><given-names>Shaozi</given-names> <surname>Li</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Yi</given-names> <surname>Yang</surname></string-name></person-group>. <article-title>&#x201C;Random Erasing Data Augmentation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the AAAI Conference on Artificial Intelligence (AAAI)</italic></source>. <year>2020</year> (<comment>cit. on pp. 111, 129, 142, 176</comment>).</mixed-citation></ref>
<ref id="CIT453"><label>[Zho+16]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Bolei</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Aditya</given-names> <surname>Khosla</surname></string-name>, <string-name><given-names>Agata</given-names> <surname>Lapedriza</surname></string-name>, <string-name><given-names>Aude</given-names> <surname>Oliva</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Antonio</given-names> <surname>Torralba</surname></string-name></person-group>. <article-title>&#x201C;Learning Deep Features for Discriminative Localization&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2016</year>, <comment>pp.</comment> <fpage>2921</fpage>&#x2013;<lpage>2929</lpage> (<comment>cit. on p. 164</comment>).</mixed-citation></ref>
<ref id="CIT454"><label>[Zho+22]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xilong</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name>, <string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name>, <string-name><given-names>Paul</given-names> <surname>Guerrero</surname></string-name>, <string-name><given-names>Kalyan</given-names> <surname>Sunkavalli</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Nima</given-names> <surname>Kalantari</surname></string-name></person-group>. <article-title>&#x201C;TileGen: Tileable, Controllable Material Generation and Capture&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (Proc. SIGGRAPH Asia)</italic></source> (<year>2022</year>) (<comment>cit. on pp. 105, 107, 109, 116, 123</comment>).</mixed-citation></ref>
<ref id="CIT455"><label>[Zho+23]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xilong</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Milo&#x0161;</given-names> <surname>Ha&#x0161;an</surname></string-name>, <string-name><given-names>Valentin</given-names> <surname>Deschaintre</surname></string-name>, <string-name><given-names>Paul</given-names> <surname>Guerrero</surname></string-name>, <string-name><given-names>Kalyan</given-names> <surname>Sunkavalli</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Nima</given-names> <surname>Khademi Kalantari</surname></string-name></person-group>. <article-title>&#x201C;A Semi-Procedural Convolutional Material Prior&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <publisher-name>Wiley Online Library</publisher-name>. <year>2023</year> (<comment>cit. on p. 105</comment>).</mixed-citation></ref>
<ref id="CIT456"><label>[ZK21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Xilong</given-names> <surname>Zhou</surname></string-name></person-group> <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Nima</given-names> <surname>Khademi Kalantari</surname></string-name></person-group>. <article-title>&#x201C;Adversarial Single-Image SVBRDF Estimation with Hybrid Training&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 40</volume>. <issue>2</issue>. <publisher-name>Wiley Online Library</publisher-name>. <year>2021</year>, <comment>pp.</comment> <fpage>315</fpage>&#x2013;<lpage>325</lpage> (<comment>cit. on pp. 123, 125, 127, 137, 138, 141, 152, 153</comment>).</mixed-citation></ref>
<ref id="CIT457"><label>[Zho+17]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yang</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Huajie</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>Dani</given-names> <surname>Lischinski</surname></string-name>, <string-name><given-names>Minglun</given-names> <surname>Gong</surname></string-name>, <string-name><given-names>Johannes</given-names> <surname>Kopf</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hui</given-names> <surname>Huang</surname></string-name></person-group>. <article-title>&#x201C;Analysis and Controlled <target target-type="page" id="pges_267"/>Synthesis of Inhomogeneous Textures&#x201D;</article-title>. <comment>In</comment>: <source><italic>Computer Graphics Forum</italic></source>. <volume>Vol. 36</volume>. <issue>2</issue>. <year>2017</year>, <comment>pp.</comment> <fpage>199</fpage>&#x2013;<lpage>212</lpage> (<comment>cit. on p. 74</comment>).</mixed-citation></ref>
<ref id="CIT458"><label>[Zho+18]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yang</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Zhen</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Xiang</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>Dani</given-names> <surname>Lischinski</surname></string-name>, <string-name><given-names>Daniel</given-names> <surname>Cohen-Or</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Hui</given-names> <surname>Huang</surname></string-name></person-group>. <article-title>&#x201C;Non-Stationary Texture Synthesis by Adversarial Expansion&#x201D;</article-title>. <comment>In</comment>: <source><italic>ACM Transactions on Graphics (TOG)</italic></source> <volume>37</volume>.<issue>4</issue> (<month>July</month> <year>2018</year>) (<comment>cit. on pp. 34, 38, 56, 72, 74, 76, 77, 79, 84, 85, 87, 90, 94, 104</comment>).</mixed-citation></ref>
<ref id="CIT459"><label>[Zho+19]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Yizhou</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Xiaoyan</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>Zheng-Jun</given-names> <surname>Zha</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Wenjun</given-names> <surname>Zeng</surname></string-name></person-group>. <article-title>&#x201C;Context-Reinforced Semantic Segmentation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>. <year>2019</year>, <comment>pp.</comment> <fpage>4046</fpage>&#x2013;<lpage>4055</lpage> (<comment>cit. on p. 41</comment>).</mixed-citation></ref>
<ref id="CIT460"><label>[Zhu+17a]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jun Yan</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Taesung</given-names> <surname>Park</surname></string-name>, <string-name><given-names>Phillip</given-names> <surname>Isola</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Alexei A.</given-names> <surname>Efros</surname></string-name></person-group>. <article-title>&#x201C;Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks&#x201D;</article-title>. <comment>In</comment>: <source><italic>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</italic></source> <comment>2017-Octob</comment> (<month>Mar.</month> <year>2017</year>), <comment>pp.</comment> <fpage>2242</fpage>&#x2013;<lpage>2251</lpage> (<comment>cit. on pp. 78, 79, 83, 94</comment>).</mixed-citation></ref>
<ref id="CIT461"><label>[Zhu+17b]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Jun-Yan</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Richard</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Deepak</given-names> <surname>Pathak</surname></string-name>, <string-name><given-names>Trevor</given-names> <surname>Darrell</surname></string-name>, <string-name><given-names>Alexei A</given-names> <surname>Efros</surname></string-name>, <string-name><given-names>Oliver</given-names> <surname>Wang</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Eli</given-names> <surname>Shechtman</surname></string-name></person-group>. <article-title>&#x201C;Toward Multimodal Image-to-Image Translation&#x201D;</article-title>. <comment>In</comment>: <source><italic>Advances in Neural Information Processing Systems</italic></source>. <year>2017</year>, <comment>pp.</comment> <fpage>465</fpage>&#x2013;<lpage>476</lpage> (<comment>cit. on p. 37</comment>).</mixed-citation></ref>
<ref id="CIT462"><label>[Zhu+21]</label> <mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Wei</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>Xian</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Dai</given-names> <surname>Owaki</surname></string-name>, <string-name><given-names>Kyo</given-names> <surname>Kutsuzawa</surname></string-name></person-group>, <comment>and</comment> <person-group person-group-type="author"><string-name><given-names>Mitsuhiro</given-names> <surname>Hayashibe</surname></string-name></person-group>. <article-title>&#x201C;A Survey of Sim-to-real Transfer Techniques Applied to Reinforcement Learning for Bioinspired Robots&#x201D;</article-title>. <comment>In</comment>: <source><italic>IEEE Transactions on Neural Networks and Learning Systems</italic></source> (<year>2021</year>) (<comment>cit. on pp. 156, 212</comment>).<target target-type="page" id="pges_268"/></mixed-citation></ref>
</ref-list>
</back>
</book-part>
<book-part id="c10-app1" book-part-type="appendix">
<book-part-meta>
<book-part-id book-part-id-type="publisher-id">URJC</book-part-id>
<title-group>
<label><target target-type="page" id="pges_269"/>A.</label>
<title>Resumen</title>
</title-group>
</book-part-meta>
<body>
<p>Las escenas virtuales hiperrealistas est&#x00E1;n cada vez m&#x00E1;s presentes en nuestra sociedad, con una amplia variedad de aplicaciones en &#x00E1;reas como la fabricaci&#x00F3;n, la arquitectura, el dise&#x00F1;o de moda y el entretenimiento, incluyendo pel&#x00ED;culas, videojuegos y realidad aumentada y virtual. Generar im&#x00E1;genes realistas de tales escenas requiere modelos de iluminaci&#x00F3;n, geometr&#x00ED;a y materiales altamente precisos, lo que puede ser costoso de obtener en recursos y tiempo. Tradicionalmente, estos modelos a menudo han sido creados manualmente por artistas entrenados, pero este proceso puede ser prohibitivamente ineficiente. Alternativamente, se pueden capturar ejemplos del mundo real, pero este enfoque presenta desaf&#x00ED;os adicionales en t&#x00E9;rminos de precisi&#x00F3;n y escalabilidad. Adem&#x00E1;s, aunque el realismo y la precisi&#x00F3;n son cruciales en esos procesos, la eficiencia en el renderizado tambi&#x00E9;n es un requisito clave, para que se puedan generar im&#x00E1;genes realistas con la velocidad requerida en muchas aplicaciones del mundo real. Uno de los desaf&#x00ED;os m&#x00E1;s significativos en este sentido es la adquisici&#x00F3;n y representaci&#x00F3;n de materiales, que son un componente cr&#x00ED;tico de nuestro mundo visual y, por extensi&#x00F3;n, de las representaciones virtuales del mismo. Sin embargo, los enfoques existentes para la adquisici&#x00F3;n y representaci&#x00F3;n de materiales est&#x00E1;n limitados en t&#x00E9;rminos de eficiencia y precisi&#x00F3;n, lo que restringe su impacto en el mundo real. Para abordar estos desaf&#x00ED;os, los enfoques basados en datos que aprovechan el aprendizaje autom&#x00E1;tico pueden proporcionar soluciones viables. Sin embargo, el dise&#x00F1;o y entrenamiento de modelos de redes neuronales que satisfagan todos estos requisitos sigue siendo un desaf&#x00ED;o, que requiere una cuidadosa consideraci&#x00F3;n de los compromisos entre calidad y eficiencia.</p>
<p>En esta tesis, proponemos soluciones novedosas basadas en aprendizaje autom&#x00E1;tico para abordar varios desaf&#x00ED;os clave en la renderizaci&#x00F3;n de escenas virtuales. Nuestros m&#x00E9;todos hacen uso de diversas formas de redes neuronales para introducir algoritmos innovadores para la codificaci&#x00F3;n de la radiancia, la generaci&#x00F3;n digital de materiales, su edici&#x00F3;n y estimaci&#x00F3;n. En primer lugar, presentamos un sistema de transferencia de atributos visuales que puede generalizar eficazmente a nuevas condiciones de iluminaci&#x00F3;n y distorsiones geom&#x00E9;tricas, que ampliamos para la adquisici&#x00F3;n de materiales de alta resoluci&#x00F3;n mediante un dispositivo dise&#x00F1;ado para tal prop&#x00F3;sito. Tambi&#x00E9;n, proponemos un modelo generativo capaz de sintetizar texturas teselables &#x015B;que se puede concatenar consigo misma sin que aparezcan artefactos de repetici&#x00F3;n&#x015B; a partir de una sola imagen de entrada, lo que ayuda a mejorar la calidad de la renderizaci&#x00F3;n de materiales. Esta tesis tambi&#x00E9;n presenta una representaci&#x00F3;n de materiales que codifica con precisi&#x00F3;n la reflectancia del material y ofrece potentes capacidades de edici&#x00F3;n y propagaci&#x00F3;n. Para lograr este objetivo, nos basamos en trabajos recientes en campos neuronales. Adem&#x00E1;s de la reflectancia, presentamos un m&#x00E9;todo novedoso para la codificaci&#x00F3;n de iluminaci&#x00F3;n <target target-type="page" id="pges_270"/>global que aprovecha modelos generativos cuidadosamente dise&#x00F1;ados para lograr un muestreo significativamente m&#x00E1;s r&#x00E1;pido y de mayor calidad que los trabajos disponibles en la literatura. Por &#x00FA;ltimo, proponemos dos m&#x00E9;todos innovadores para la digitalizaci&#x00F3;n escalable de materiales. Usando esc&#x00E1;neres como dispositivo de captura, presentamos un modelo generativo que puede proporcionar estimaciones de reflectancia de materiales de alta resoluci&#x00F3;n utilizando una sola imagen como entrada. Adicionalmente, introducimos el primer algoritmo de cuantificaci&#x00F3;n de incertidumbre que aumenta su fiabilidad y eficiencia. Adem&#x00E1;s, presentamos un m&#x00E9;todo neuronal para digitalizar las propiedades mec&#x00E1;nicas de tejidos utilizando im&#x00E1;genes de profundidad como entrada, que extendemos con una m&#x00E9;trica de similitud perceptualmente validada mediante un estudio de usuario. Globalmente, las contribuciones de esta tesis representan avances significativos en los campos de la codificaci&#x00F3;n de la radiancia y la adquisici&#x00F3;n y edici&#x00F3;n digital de materiales, lo que permite mejorar la calidad, escalabilidad y eficiencia de los sistemas de renderizaci&#x00F3;n.</p>
<sec id="c10-app1.s1">
<label>A.1</label>
<title>Antecedentes</title>
<sec id="c10-app1-s1-1">
<title>Generaci&#x00F3;n Digital de Materiales</title>
<p>La digitalizaci&#x00F3;n precisa y de alta resoluci&#x00F3;n de la reflectancia de los materiales es un problema crucial para generar entornos virtuales atractivos y realistas. Este proceso de digitalizaci&#x00F3;n se realiza t&#x00ED;picamente utilizando un <italic>goniorreflect&#x00F3;metro</italic>, una m&#x00E1;quina altamente compleja dise&#x00F1;ada para medir la respuesta fotom&#x00E9;trica de un material utilizando fuentes de luz y c&#x00E1;maras colocadas en diferentes posiciones alrededor de su semiesfera. Los dispositivos existentes para medir la apariencia del material en muestras espacialmente variables se limitan a una sola escala, ya sea micro o mesosc&#x00F3;pica. Esta es una limitaci&#x00F3;n pr&#x00E1;ctica cuando el material tiene una estructura compleja a m&#x00FA;ltiples escalas, como es com&#x00FA;n en muchos materiales del mundo real. Para ellos, medir con precisi&#x00F3;n su reflectancia espacialmente variable requiere im&#x00E1;genes microsc&#x00F3;picas de alta resoluci&#x00F3;n. Por ejemplo, el comportamiento &#x00F3;ptico de los textiles depende mucho de sus fibras o microestructuras [<xref ref-type="bibr" rid="CIT040">CLA19</xref>]. Adem&#x00E1;s, muchos materiales tambi&#x00E9;n presentan variaciones <italic>mesosc&#x00F3;picas</italic> a gran escala, como patrones de impresi&#x00F3;n o los tartanes, que no se pueden capturar utilizando &#x00FA;nicamente fotograf&#x00ED;as microsc&#x00F3;picas. Adem&#x00E1;s, muchos de estos dispositivos no pueden capturar propiedades &#x00F3;pticas importantes, como la transmitancia del material, que es clave para la representaci&#x00F3;n realista de objetos como los tejidos. No hay dispositivos de captura ni algoritmos disponibles para la digitalizaci&#x00F3;n precisa y de alta resoluci&#x00F3;n de <italic>Spatially Varying Bidirectional Scattering Distribution Function</italic> (SVBSDF) para materiales heterog&#x00E9;neos. Para introducir tales materiales, que son muy comunes en muchas industrias como la moda, la textil o la fabricaci&#x00F3;n de cuero, en los sistemas de representaci&#x00F3;n, es necesario <target target-type="page" id="pges_271"/>crear estas m&#x00E1;quinas y desarrollar algoritmos precisos para manejar los datos que capturan. Finalmente, dicho dispositivo puede servir como fuente de generaci&#x00F3;n de datos para entrenar modelos que operen en escenarios de datos m&#x00E1;s limitados, como en la digitalizaci&#x00F3;n a partir de una sola imagen.</p>
</sec>
<sec id="c10-app1-s1-2">
<title>Edici&#x00F3;n y S&#x00ED;ntesis</title>
<p><bold>SVBRDFs Teselables</bold> Incluso si la reflectancia puede ser medida con precisi&#x00F3;n para una gran superficie de un material, en entornos de renderizaci&#x00F3;n, tambi&#x00E9;n es importante que el material digital sea <italic>teselable</italic>. Una textura teselable es una imagen de textura que puede ser repetida infinitamente en todas las direcciones sin presentar discontinuidades visibles. En otras palabras, una textura teselable est&#x00E1; dise&#x00F1;ada de tal manera que se conecta perfectamente con copias id&#x00E9;nticas de s&#x00ED; misma, que pueden ser dispuestas una al lado de la otra o apiladas una encima de la otra para crear un patr&#x00F3;n m&#x00E1;s grande y continuo sin ninguna repetici&#x00F3;n visible. Esto se utiliza a menudo en gr&#x00E1;ficos por ordenador, especialmente para crear texturas en modelos 3D, o incluso para la colocaci&#x00F3;n de fondos en sitios web. En los &#x00FA;ltimos a&#x00F1;os, se han propuesto diferentes m&#x00E9;todos para la generaci&#x00F3;n de texturas teselables a partir de un solo ejemplo de entrada [<xref ref-type="bibr" rid="CIT281">Mor+17</xref>; <xref ref-type="bibr" rid="CIT341">Rod+19</xref>; <xref ref-type="bibr" rid="CIT279">Mor+20</xref>]; sin embargo, estos trabajos presentan varias limitaciones para la generaci&#x00F3;n de <italic>Spatially Varying Bidirectional Reflectance Distribution Functions</italic> (SVBRDFs) teselables. En primer lugar, todos suponen un cierto nivel de regularidad en las im&#x00E1;genes de entrada, o bien las texturas generadas pierden una cantidad significativa de fidelidad visual. Adem&#x00E1;s, estos m&#x00E9;todos se limitan a generar un solo mapa de textura (normalmente, el albedo del material). Esto es problem&#x00E1;tico para la s&#x00ED;ntesis de SVBRDFs teselables, que requieren transformar cada mapa de textura en el SVBRDF en un mapa teselable, al mismo tiempo que se preserva la coherencia p&#x00ED;xel a p&#x00ED;xel entre los mapas. En resumen, en la actualidad no es posible generar SVBRDFs teselables de alta calidad a partir de un solo ejemplo de entrada con los m&#x00E9;todos propuestos en trabajos previos.</p>
<p><bold>Transferencia de Atributos Visuales</bold> Las representaciones efectivas de materiales requieren comprender las propiedades que los definen de manera &#x00FA;nica, a las que nos referimos como <italic>atributos visuales</italic>. Estos atributos son par&#x00E1;metros que var&#x00ED;an espacialmente y mantienen la coherencia espacial con respecto a la estructura del material, mientras que permanecen invariantes ante los cambios en la iluminaci&#x00F3;n de la escena o la geometr&#x00ED;a del objeto subyacente. Por ejemplo, pueden representar propiedades &#x00F3;pticas de un SVBRDF, estilizaciones art&#x00ED;sticas o m&#x00E1;scaras de segmentaci&#x00F3;n sem&#x00E1;ntica. Obtener estos atributos para muestras grandes de materiales es un problema que presenta numerosas dificultades. Una forma de abordarlo es obtener estos atributos en un peque&#x00F1;o ejemplo del material y <italic>transferirlos</italic> a porciones m&#x00E1;s grandes del mismo, utilizando una imagen de entrada grande que sirve como gu&#x00ED;a. En los &#x00FA;ltimos a&#x00F1;os, se han propuesto diferentes m&#x00E9;todos <target target-type="page" id="pges_272"/>de transferencia de atributos visuales, sin embargo, est&#x00E1;n limitados en su eficiencia, robustez, generalizaci&#x00F3;n o rendimiento. Idealmente, un m&#x00E9;todo de transferencia de atributos visuales para materiales deber&#x00ED;a ser robusto ante nuevas condiciones de iluminaci&#x00F3;n del material, degradaciones de la c&#x00E1;mara o distorsiones de entrada; adem&#x00E1;s de ser eficiente (proporcionando tiempos de transferencia interactivos incluso en presupuestos computacionales bajos), controlable y predecible, y capaz de generar salidas de alta resoluci&#x00F3;n, que son necesarias para SVBRDFs realistas. Desarrollar un m&#x00E9;todo con estas caracter&#x00ED;sticas puede resultar &#x00FA;til para muchas tareas posteriores, como sistemas eficientes de generaci&#x00F3;n de datos o digitalizaci&#x00F3;n de materiales.</p>
</sec>
<sec id="c10-app1-s1-3">
<title>Digitalizaci&#x00F3;n Escalable</title>
<p><bold>Estimaci&#x00F3;n de la Reflectancia de Materiales</bold> Como ya hemos mencionado, la digitalizaci&#x00F3;n precisa de los materiales t&#x00ED;picamente requiere dispositivos costosos que utilizan cantidades significativas de tiempo e intervenci&#x00F3;n manual para operar. Estos dispositivos proporcionan representaciones digitales altamente realistas de la reflectancia del material a costa de la escalabilidad, lo que limita su aplicabilidad en muchos entornos del mundo real. Por ejemplo, muchos procesos de fabricaci&#x00F3;n, como los de la industria textil, generan una amplia variedad de diferentes materiales que no pueden ser digitalizados f&#x00E1;cil ni econ&#x00F3;micamente con la cadencia requerida en estas industrias. Adem&#x00E1;s, la disponibilidad limitada de estos dispositivos introduce ineficiencias adicionales y costes econ&#x00F3;micos y ecol&#x00F3;gicos, ya que las muestras f&#x00ED;sicas de los materiales deben ser enviadas desde el cliente a la localizaci&#x00F3;n del dispositivo de captura.</p>
<p>La estimaci&#x00F3;n de la reflectancia de materiales a partir de dispositivos de bajo coste es un problema que se ha estudiado ampliamente en la literatura [<xref ref-type="bibr" rid="CIT004">AWL13</xref>; <xref ref-type="bibr" rid="CIT005">AWL15</xref>; <xref ref-type="bibr" rid="CIT065">Des+18</xref>; <xref ref-type="bibr" rid="CIT148">Hen+21</xref>; <xref ref-type="bibr" rid="CIT127">Guo+20b</xref>; <xref ref-type="bibr" rid="CIT400">VPS21</xref>]. Estos m&#x00E9;todos prometen proporcionar una digitalizaci&#x00F3;n de materiales escalable, eficiente y econ&#x00F3;mica. Normalmente, se basan en una o varias im&#x00E1;genes de un material iluminadas con flash tomadas con una c&#x00E1;mara de smartphone, y en un modelo de aprendizaje profundo entrenado en materiales sint&#x00E9;ticos que proporciona estimaciones plausibles. Sin embargo, estos enfoques presentan varias limitaciones que los hacen inadecuados para flujos de trabajo en los que se requiere digitalizaci&#x00F3;n precisa. En primer lugar, muchos de estos m&#x00E9;todos se basan en modelos generativos que tienden a producir artefactos o utilizan grandes redes neuronales que limitan la resoluci&#x00F3;n de salida, lo que reduce la calidad de los materiales estimados. Adem&#x00E1;s, las im&#x00E1;genes tomadas con smartphones presentan importantes desaf&#x00ED;os de calibraci&#x00F3;n, mientras que la mayor&#x00ED;a de los conjuntos de datos existentes utilizados para entrenar estos modelos son puramente sint&#x00E9;ticos, lo que limita a&#x00FA;n m&#x00E1;s la calidad de estas estimaciones. Adem&#x00E1;s, muchos de estos m&#x00E9;todos utilizan funciones de coste perceptuales para el entrenamiento o la optimizaci&#x00F3;n en tiempo de evaluaci&#x00F3;n, <target target-type="page" id="pges_273"/>lo que, si bien ayuda a digitalizar materiales estoc&#x00E1;sticos, presenta desaf&#x00ED;os adicionales para garantizar la repetibilidad y consistencia necesarias para construir un inventario digital. Adem&#x00E1;s de estas limitaciones, la estimaci&#x00F3;n de reflectancia de materiales a partir de una sola imagen sigue siendo un problema mal planteado y, por lo tanto, muchos resultados diferentes pueden ser plausibles dados los mismos datos de entrada. Hasta ahora, no existe una forma de cuantificar la incertidumbre para este problema, lo que dificulta la aplicabilidad de estos sistemas en escenarios del mundo real, y limita su fiabilidad y eficiencia.</p>
<p><bold>Comportamiento Mec&#x00E1;nico de Materiales Textiles</bold> Adem&#x00E1;s de la reflectancia de los materiales, otro componente importante de las escenas virtuales es la geometr&#x00ED;a de los objetos. Mientras que muchos objetos del mundo real tienen formas est&#x00E1;ticas, que pueden ser escaneadas con dispositivos relativamente accesibles como el sensor <italic>LiDAR</italic> en <italic>iPhones</italic>, existen materiales que cambian de forma cuando se les aplican fuerzas externas, como la gravedad. En este sentido, los materiales deformables requieren modelos computacionales y sistemas de captura espec&#x00ED;ficos. Uno de los materiales m&#x00E1;s comunes de este tipo son las telas, que son ubicuas en el mundo real y se utilizan en nuestra ropa, muebles o veh&#x00ED;culos. Para ellas, capturar con precisi&#x00F3;n su forma est&#x00E1;tica no es suficiente: Es necesario poder simular su comportamiento en nuevos escenarios, como nuevas prendas o fuerzas din&#x00E1;micas, como el viento [<xref ref-type="bibr" rid="CIT029">Ber+23</xref>], o el movimiento humano [<xref ref-type="bibr" rid="CIT352">SOC22</xref>]. Adem&#x00E1;s, las telas son incre&#x00ED;blemente variadas en sus patrones de fabricaci&#x00F3;n (por ejemplo, diferentes patrones de tejido o estructuras de punto), composiciones y acabados, lo que no solo determina su apariencia &#x00F3;ptica [<xref ref-type="bibr" rid="CIT040">CLA19</xref>] sino tambi&#x00E9;n su comportamiento mec&#x00E1;nico. El problema de capturar las propiedades mec&#x00E1;nicas de las telas para que puedan ser re-simuladas en entornos virtuales ha sido ampliamente estudiado en la literatura; sin embargo, las soluciones actuales requieren una intervenci&#x00F3;n humana tediosa o dispositivos de captura costosos, lo que dificulta su escalabilidad. Adem&#x00E1;s, existe una falta de entendimiento sobre la similitud perceptual del comportamiento mec&#x00E1;nico de las telas, lo que crea desaf&#x00ED;os adicionales para dise&#x00F1;ar sistemas precisos de estimaci&#x00F3;n de par&#x00E1;metros mec&#x00E1;nicos de tejidos.</p>
</sec>
<sec id="c10-app1-s1-4">
<title>Codificaci&#x00F3;n Neuronal de Radiancia</title>
<p><bold>Reflectancia</bold> Las <italic>Bidirectional Texture Functions (BTFs)</italic> proporcionan representaciones densas y de gran exactitud de la reflectancia de los materiales, sin embargo, son t&#x00ED;picamente prohibitivas en t&#x00E9;rminos computacionales y de memoria. Trabajo reciente [<xref ref-type="bibr" rid="CIT317">Rai+19</xref>; <xref ref-type="bibr" rid="CIT316">Rai+20</xref>; <xref ref-type="bibr" rid="CIT220">Kuz+21</xref>] presenta representaciones neuronales para BTFs, consiguiendo resultados de muy alta calidad y eficiencia. Los m&#x00E9;todos neuronales de codificaci&#x00F3;n de reflectancia entrenados en BTF sint&#x00E9;ticos o medidos han mostrado resultados prometedores para aumentar el realismo de los materiales renderizados. Sin embargo, estas <target target-type="page" id="pges_274"/>codificaciones son inmutables, lo que significa que su salida para una consulta determinada de UV, c&#x00E1;mara y vector de luz queda fija una vez que se entrenan. Esto puede ser muy limitante cuando el fragmento del material utilizado para el entrenamiento es demasiado peque&#x00F1;o o no se puede teselar, lo que ocurre</p>
<p>con frecuencia cuando el material ha sido escaneado con un dispositivo de captura. Por lo tanto, las codificaciones de reflectancia neuronal actuales, aunque precisas y eficientes, tienen importantes limitaciones en cuanto a las capacidades de edici&#x00F3;n, lo que dificulta su aplicabilidad y utilidad en entornos de renderizado del mundo real.</p>
<p><bold>Iluminaci&#x00F3;n</bold> Existen numerosas aproximaciones para la representaci&#x00F3;n de iluminaci&#x00F3;n en escenas virtuales, incluyendo <italic>Spherical Harmonics</italic>, <italic>Spherical Gaussians</italic> o <italic>Haar Wavelets</italic>. Estas, sin embargo, tienen dificultades para representar de manera eficiente y precisa condiciones de iluminaci&#x00F3;n no difusas, como fuentes de luz de alta frecuencia, como el sol o las bombillas. Trabajo reciente sobre representaciones neuronales de la iluminaci&#x00F3;n natural [<xref ref-type="bibr" rid="CIT105">GES22</xref>] proporciona mejores aproximaciones para estos casos, pero tiene dificultades con la iluminaci&#x00F3;n interior o las escenas nocturnas con m&#x00FA;ltiples fuentes de luz. Adem&#x00E1;s, estas aproximaciones neuronales no proporcionan capacidades de muestreo y evaluaci&#x00F3;n de PDF, que son esenciales para el renderizado de Monte Carlo en el trazado de rayos. No existen m&#x00E9;todos de aproximaci&#x00F3;n de iluminaci&#x00F3;n que proporcionen estas capacidades de muestreo, a la vez que funcionen con precisi&#x00F3;n y eficiencia en cualquier tipo de mapa de entorno de entrada. Trabajo reciente sobre modelos generativos invertibles y representaciones neuronales impl&#x00ED;citas puede proporcionar un camino prometedor para resolver estos problemas.</p>
</sec>
</sec>
<sec id="c10-app1.s2">
<label>A.2</label>
<title>Objetivos</title>
<p>El objetivo principal de esta tesis es crear algoritmos innovadores para la representaci&#x00F3;n computacional y la captura de materiales y la radiancia de escenas. Esto es importante por varios motivos. Estas representaciones digitales pueden ser utilizadas en una amplia variedad de aplicaciones y problemas, incluyendo gr&#x00E1;ficos por ordenador, videojuegos, producci&#x00F3;n de pel&#x00ED;culas, realidad virtual y aumentada, o dise&#x00F1;o computacional. Tienen el potencial de aumentar significativamente el realismo visual y la calidad de los entornos virtuales creados en estos sistemas. Adem&#x00E1;s, la precisi&#x00F3;n y la eficiencia de estos m&#x00E9;todos tienen un impacto directo en la calidad y los recursos computacionales requeridos para estas tareas. Por lo tanto, obtener representaciones y algoritmos de digitalizaci&#x00F3;n precisos y eficientes para materiales y otros componentes de las escenas virtuales puede ayudar a reducir el costo computacional de la renderizaci&#x00F3;n, el dise&#x00F1;o, la captura y la modelizaci&#x00F3;n, lo que puede llevar a sistemas m&#x00E1;s r&#x00E1;pidos, <target target-type="page" id="pges_275"/>eficientes y escalables. Adem&#x00E1;s, pueden permitir la creaci&#x00F3;n procedural eficiente de conjuntos de datos para entrenar modelos para resolver otras tareas, como segmentaci&#x00F3;n o clasificaci&#x00F3;n, o generaci&#x00F3;n de im&#x00E1;genes o videos. Finalmente, adem&#x00E1;s de la eficiencia y la precisi&#x00F3;n, caracter&#x00ED;sticas adicionales como la diferenciabilidad o la capacidad de cuantificaci&#x00F3;n de la incertidumbre crean oportunidades para aplicaciones de optimizaci&#x00F3;n o el aprendizaje activo.</p>
<p>En esta tesis, presentamos nuevos m&#x00E9;todos basados en aprendizaje autom&#x00E1;tico para abordar los problemas descritos anteriormente. Nuestras soluciones propuestas buscan lograr la m&#x00E1;xima eficiencia computacional y de datos, as&#x00ED; como la precisi&#x00F3;n, el control y la fiabilidad. Adem&#x00E1;s, nuestros algoritmos est&#x00E1;n dise&#x00F1;ados para ser beneficiosos para los usuarios finales al incorporar requisitos humanos como componentes perceptuales, de edici&#x00F3;n, predictibilidad, robustez y cuantificaci&#x00F3;n de incertidumbre. Todos nuestros m&#x00E9;todos propuestos aprovechan las capacidades de las redes neuronales, incorporando los &#x00FA;ltimos avances en representaciones impl&#x00ED;citas, dise&#x00F1;o de arquitectura y funciones de coste, y procedimientos de entrenamiento. Todos nuestros modelos son completamente diferenciables y se pueden incorporar en los procesos de optimizaci&#x00F3;n para resolver problemas de renderizaci&#x00F3;n inversa. Nuestras soluciones para la codificaci&#x00F3;n de la radiancia se pueden incorporar sin problemas en motores de trazado de rayos para aumentar la eficiencia del renderizado o en representaciones neuronales de escenas para aumentar el realismo. Adem&#x00E1;s, nuestros m&#x00E9;todos de digitalizaci&#x00F3;n casual de alta calidad permiten la captura escalable de materiales, lo que puede ayudar a los usuarios finales a crear sus propios materiales virtuales realistas o crear grandes conjuntos de datos a un menor coste.</p>
<p>Los objetivos de esta tesis se pueden resumir en tres puntos:</p>
<list list-type="bullet">
<list-item><p>Dise&#x00F1;ar nuevos m&#x00E9;todos basados en algoritmos de aprendizaje autom&#x00E1;tico para la generaci&#x00F3;n, propagaci&#x00F3;n, edici&#x00F3;n y s&#x00ED;ntesis de materiales digitales, con altos est&#x00E1;ndares de predecibilidad, control y calidad.</p></list-item>
<list-item><p>Introducir nuevas representaciones para la radiancia de escenas, que sean eficientes, editables y diferenciables.</p></list-item>
<list-item><p>Democratizar la captura digital de alta calidad de materiales mediante el dise&#x00F1;o de nuevos m&#x00E9;todos de bajo coste, escalables, fiables y validados perceptualmente.</p></list-item>
</list>
</sec>
<sec id="c10-app1.s3">
<label>A.3</label>
<title>Metodolog&#x00ED;a</title>
<p>En esta secci&#x00F3;n detallamos la metodolog&#x00ED;a seguida durante el transcurso de esta tesis.</p>
<sec id="c10-app1-s3-1">
<title><target target-type="page" id="pges_276"/>Revisi&#x00F3;n Bibliogr&#x00E1;fica</title>
<p>Una vez definidos los objetivos de la tesis, descritos anteriormente, se revis&#x00F3; la bibliograf&#x00ED;a de trabajos previos para identificar huecos en la literatura y encontrar l&#x00ED;neas de investigaci&#x00F3;n prometedoras. En este proceso de revisi&#x00F3;n de la bibliograf&#x00ED;a, encontramos una importante ausencia de m&#x00E9;todos capaces de generar datos de alta calidad para el entrenamiento de modelos de captura y representaci&#x00F3;n de materiales; adem&#x00E1;s de la oportunidad de entrenar modelos de captura y edici&#x00F3;n de materiales m&#x00E1;s precisos, eficientes y robustos. Es importante recalcar que esta revisi&#x00F3;n se realiz&#x00F3; durante todas las etapas de la tesis, ya que los campos del aprendizaje autom&#x00E1;tico, la inform&#x00E1;tica gr&#x00E1;fica y la visi&#x00F3;n por ordenador se mueven a una velocidad muy elevada. En ellos, nuevas soluciones innovadoras se proponen a diario, por lo que mantener vivo este an&#x00E1;lisis del estado del arte es fundamental para el desarrollo efectivo de algoritmos. Por ejemplo, nuevas tendencias como los campos neuronales [<xref ref-type="bibr" rid="CIT276">Mil+20</xref>] y las representaciones impl&#x00ED;citas [<xref ref-type="bibr" rid="CIT366">Sit+20</xref>] han sido identificadas y han sido claves para el desarrollo de muchos de los algoritmos propuestos en esta tesis.</p>
</sec>
<sec id="c10-app1-s3-2">
<title>Generaci&#x00F3;n de Materiales Digitales</title>
<p>Los primeros proyectos desarrollados en el contexto de esta tesis est&#x00E1;n enfocados en la generaci&#x00F3;n materiales digitales de alta calidad, los cuales pueden ser utilizados posteriormente para otras tareas, como entrenar representaciones neuronales de materiales o crear grandes conjuntos de datos para modelos de estimaci&#x00F3;n de materiales. Estos m&#x00E9;todos toman como entrada un solo material, lo que ayuda a aumentar la predictibilidad de sus resultados.</p>
<p>En el <xref ref-type="book-part" rid="c2">cap&#x00ED;tulo 2</xref>, presentamos un m&#x00E9;todo neuronal para la de propagaci&#x00F3;n de atributos visuales capaz de transferir, para un material dado, muchos tipos de mapas de propiedades visuales a im&#x00E1;genes de nuevos parches del mismo material, o uno similar. Estas nuevas im&#x00E1;genes pueden ser capturadas bajo cualquier condici&#x00F3;n de iluminaci&#x00F3;n, variaciones de color, distorsiones geom&#x00E9;tricas o configuraci&#x00F3;n de c&#x00E1;mara. Este modelo utiliza redes neuronales convolucionales de imagen a imagen con pocos par&#x00E1;metros, para proporcionar la primera soluci&#x00F3;n capaz de aprovechar el comportamiento &#x00F3;ptico del material para el problema de transferencia de atributos visuales, al ser entrenado con un conjunto de datos fotom&#x00E9;trico y una pol&#x00ED;tica amplia de aumentado sint&#x00E9;tico de datos. Este m&#x00E9;todo puede ser entrenado en menos de un minuto y es capaz de transferir cualquier tipo de atributo visual, desde las normales de la superficie, hasta estilizaciones art&#x00ED;sticas o mapas de segmentaci&#x00F3;n sem&#x00E1;ntica. Sin embargo, est&#x00E1; limitado en precisi&#x00F3;n, s&#x00F3;lo puede aprender de un conjunto de datos a la vez y s&#x00F3;lo puede aprender a transferir un atributo visual.</p>
<p><target target-type="page" id="pges_277"/>Para abordar estas limitaciones, en el <xref ref-type="sec" rid="c2-sC">cap&#x00ED;tulo 2.C</xref> ampliamos este m&#x00E9;todo, presentando una red neuronal m&#x00E1;s compleja capaz de generar mapas de SVBSDF de gran calidad y alta resoluci&#x00F3;n para cualquier material, dados uno o varios conjuntos de datos fotom&#x00E9;tricos que representan peque&#x00F1;as porciones del material. A trav&#x00E9;s de extensas modificaciones en el procedimiento de entrenamiento y mejoras en la arquitectura del modelo, no s&#x00F3;lo habilitamos esta transferencia de atributos visuales con la capacidad de aprender de m&#x00FA;ltiples conjuntos</p>
<p>de datos y transferir varias im&#x00E1;genes simult&#x00E1;neamente, sino que tambi&#x00E9;n aumentamos la precisi&#x00F3;n y nitidez de sus estimaciones. Sin embargo, estas mejoras incrementan el coste computacional del entrenamiento de la red neuronal y limitan su generalidad. Con estos modelos, podemos construir materiales a gran escala, que pueden ser utilizados para generar datos BTF para entrenar otras representaciones (<xref ref-type="book-part" rid="c4">Cap&#x00ED;tulo 4</xref>), o construir grandes conjuntos de datos para entrenar modelos que se generalicen a nuevos materiales (<xref ref-type="book-part" rid="c5">Cap&#x00ED;tulo 5</xref>).</p>
<p>Las texturas utilizadas en aplicaciones de renderizaci&#x00F3;n, adem&#x00E1;s de representar una gran parte de las variaciones del material, necesitan ser teselables. Es decir, cuando se concatenan espacialmente consigo mismas, no debe ser posible identificar estas repeticiones. En el <xref ref-type="book-part" rid="c3">cap&#x00ED;tulo 3</xref>, presentamos una Red Generativa Antag&#x00F3;nica (GAN) entrenada con una sola imagen, que puede generar m&#x00FA;ltiples texturas teselables del mismo material. Para lograr esto, ampliamos trabajos anteriores sobre s&#x00ED;ntesis de texturas para permitir texturas teselables mediante un nuevo algoritmo de manipulaci&#x00F3;n de espacios latentes; y para manejar m&#x00FA;ltiples mapas de textura simult&#x00E1;neamente, mediante modificaciones en la arquitectura neuronal. Adem&#x00E1;s, aprovechamos el discriminador de la GAN como funci&#x00F3;n de estimaci&#x00F3;n de calidad en tiempo de evaluaci&#x00F3;n, lo que ayuda a distinguir entre generaciones de alta y baja calidad. Este modelo logra texturas teselables de mayor calidad que trabajos anteriores, en tiempos computacionales m&#x00E1;s bajos.</p>
<p>Estos modelos dependen de un solo material como entrada, lo que, si bien aumenta la predictibilidad, limita sus capacidades de generalizaci&#x00F3;n. No obstante, estos m&#x00E9;todos proporcionan evidencia de que, bajo ciertas suposiciones y mediante procedimientos de entrenamiento altamente controlados, las redes neuronales pueden ser utilizadas para generar datos de alta calidad. Estos datos pueden ser aprovechados para otras tareas, como se demuestra en otras secciones de esta tesis.</p>
</sec>
<sec id="c10-app1-s3-3">
<title>Codificaci&#x00F3;n Neuronal de Radiancia</title>
<p>Uno de los problemas de investigaci&#x00F3;n que exploramos en esta tesis es la representaci&#x00F3;n de radiancia utilizando redes neuronales, lo que ofrece varios <target target-type="page" id="pges_278"/>beneficios potenciales en comparaci&#x00F3;n con enfoques tradicionales. Las redes neuronales pueden proporcionar funciones eficientes, continuas y completamente diferenciables que son &#x00FA;tiles para la renderizaci&#x00F3;n directa y para problemas de renderizaci&#x00F3;n inversa. Adem&#x00E1;s, pueden permitir novedosas capacidades de edici&#x00F3;n o muestreo eficiente para la renderizaci&#x00F3;n mediante Monte Carlo. En este trabajo, nos centramos en dos componentes clave de las escenas virtuales: los materiales y la iluminaci&#x00F3;n. Nuestro objetivo es desarrollar enfoques basados en redes neuronales para codificar estos elementos que sean precisos y computacionalmente eficientes, abriendo el camino para nuevas aplicaciones en inform&#x00E1;tica gr&#x00E1;fica y visi&#x00F3;n por ordenador.</p>
<p>Respecto a los materiales en el <xref ref-type="book-part" rid="c4">cap&#x00ED;tulo 4</xref> presentamos una nueva representaci&#x00F3;n neuronal para codificar su reflectancia. Construyendo sobre trabajos previos sobre materiales neuronales [<xref ref-type="bibr" rid="CIT220">Kuz+21</xref>], nuestro m&#x00E9;todo tiene tres componentes: Una <italic>textura neuronal</italic>, que codifica la reflectancia del material en un espacio latente 2D; un <italic>renderizador neuronal</italic> completamente convolucional y con activaciones sinusoidales, que toma como entrada esta reflectancia latente y los &#x00E1;ngulos de luz y visi&#x00F3;n y genera valores lineales RGB; y un <italic>autoencoder</italic>, que puede estimar la textura neuronal de cualquier imagen de entrada 2D del material. Al igual que en trabajos previos sobre campos neuronales [<xref ref-type="bibr" rid="CIT276">Mil+20</xref>], entrenamos cada modelo desde cero para cada material, tomando como entrada su <italic>Bidirectional Texture Function</italic> (BTF), que puede ser capturada o generada sint&#x00E9;ticamente. Nuestro enfoque introduce la primera representaci&#x00F3;n neuronal para BTFs con entrada condicional, lo que hace posible la extrapolaci&#x00F3;n de BTFs, la creaci&#x00F3;n de BTF teselables y la s&#x00ED;ntesis de nuevos materiales mediante propagaci&#x00F3;n de reflectancia. Abordamos una de las principales limitaciones de los trabajos previos sobre materiales neuronales, que era su falta de capacidades de edici&#x00F3;n, lo que imped&#x00ED;a su aplicabilidad a escenarios de renderizaci&#x00F3;n del mundo real que requieren materiales de gran escala y teselables. A pesar de estas mejoras, nuestra representaci&#x00F3;n tiene limitaciones. Por ejemplo, propiedades sem&#x00E1;nticas como la especularidad o el albedo se codifican en el espacio latente y, por tanto, no son f&#x00E1;cilmente editables. Sin embargo, creemos que nuestro m&#x00E9;todo tiene un potencial significativo para trabajos futuros sobre representaciones neuronales de materiales.</p>
<p>En el <xref ref-type="book-part" rid="c7">cap&#x00ED;tulo 7</xref>, abordamos el problema de la codificaci&#x00F3;n de la iluminaci&#x00F3;n, presentando un nuevo m&#x00E9;todo neuronal para el muestreo, la codificaci&#x00F3;n de la PDF y la compresi&#x00F3;n RGB de los mapas de entorno utilizados para la iluminaci&#x00F3;n global en la renderizaci&#x00F3;n. Nuestro m&#x00E9;todo se basa en trabajos anteriores sobre <italic>representaciones neuronales impl&#x00ED;citas</italic> para comprimir el mapa de entorno HDRi en una funci&#x00F3;n continua y diferenciable. Dise&#x00F1;amos un <italic>normalizing flow</italic> que puede muestrear direcciones de mapas de entorno y medir su densidad de probabilidad. Ambas funcionalidades son esenciales para el muestreo por importancia en la renderizaci&#x00F3;n mediante Monte Carlo. Nuestra dise&#x00F1;o de arquitectura neuronal y de su procedimientos de <target target-type="page" id="pges_279"/>entrenamiento nos permitien obtener modelos que son hasta dos &#x00F3;rdenes de magnitud m&#x00E1;s r&#x00E1;pidos que los m&#x00E9;todos anal&#x00ED;ticos. Utilizando un conjunto de datos de mapas de entorno con propiedades diversas, demostramos alta generalidad y precisi&#x00F3;n, superando trabajos anteriores en este campo. Para fomentar futuras investigaciones, proporcionaremos una implementaci&#x00F3;n de c&#x00F3;digo abierto y un conjunto de datos de modelos entrenados. Si bien este es el primer m&#x00E9;todo que aprende a muestrear de mapas de entorno, creemos que existe potencial para mejorar a&#x00FA;n m&#x00E1;s nuestro enfoque mediante la introducci&#x00F3;n de dise&#x00F1;os de redes neuronales m&#x00E1;s sofisticados o aprendiendo distribuciones a priori sobre la iluminaci&#x00F3;n global.</p>
</sec>
<sec id="c10-app1-s3-4">
<title>Digitalizaci&#x00F3;n Escalable de Materiales</title>
<p>Uno de los objetivos clave de esta tesis es desarrollar soluciones de digitalizaci&#x00F3;n que sean escalables y asequibles, sin depender de hardware costoso o inaccesible. Para lograrlo, hemos realizado dos contribuciones diferentes. En primer lugar, en el <xref ref-type="book-part" rid="c5">cap&#x00ED;tulo 5</xref>, presentamos un m&#x00E9;todo novedoso para estimar las <italic>Spatially-Varying Bidirectional Reflectance Distribution Functions</italic> (SVBRDF) utilizando esc&#x00E1;neres como dispositivo de captura. Este m&#x00E9;todo se basa en gran medida en los algoritmos de generaci&#x00F3;n de datos presentados en secciones anteriores de esta tesis. Nuestro enfoque utiliza una red neuronal ligera con mecanismos de atenci&#x00F3;n, lo que resulta en estimaciones con una resoluci&#x00F3;n considerablemente mayor que los m&#x00E9;todos anteriores que se utilizaban im&#x00E1;genes tomadas con smartphones como entrada. Observamos emp&#x00ED;ricamente que la reflectancia del material depende en gran medida de su microgeometr&#x00ED;a, que los esc&#x00E1;neres planos pueden capturar adecuadamente, a la vez que proporcionan im&#x00E1;genes que se pueden utilizar directamente como el <italic>albedo</italic> de la SVBRDF. Una de las contribuciones clave de este trabajo es la introducci&#x00F3;n del primer algoritmo de cuantificaci&#x00F3;n de incertidumbre para la estimaci&#x00F3;n de materiales, para el cual demostramos aplicaciones en la creaci&#x00F3;n de conjuntos de datos a trav&#x00E9;s del aprendizaje activo. Adem&#x00E1;s, una extensi&#x00F3;n de este trabajo forma parte de <uri xlink:href="http://www.Textura.ai">Textura.ai</uri>, que es utilizada por usuarios reales, lo que demuestra su robustez y fiabilidad. Tenemos confianza en que las novedades introducidas por este m&#x00E9;todo dar&#x00E1;n forma a futuros avances sobre el problema de la digitalizaci&#x00F3;n de materiales a partir de una sola imagen. Aunque confiamos en las capacidades de nuestro m&#x00E9;todo, este se puede extender para estimar propiedades de reflectancia adicionales como la anisotrop&#x00ED;a o la transmitancia, para generalizarse a otros materiales o dispositivos, y para mejorar a&#x00FA;n m&#x00E1;s su exactitud.</p>
<p>Adem&#x00E1;s, en el <xref ref-type="book-part" rid="c6">cap&#x00ED;tulo 6</xref>, proponemos un m&#x00E9;todo para la estimaci&#x00F3;n de par&#x00E1;metros mec&#x00E1;nicos de materiales textiles utilizando im&#x00E1;genes de profundidad y la densidad del material como entradas. Nuestro m&#x00E9;todo logra una mayor escalabilidad que los trabajos anteriores, que requer&#x00ED;an secuencias de v&#x00ED;deo, <target target-type="page" id="pges_280"/>intervenci&#x00F3;n manual o dispositivos costosos y espec&#x00ED;ficos. Para lograrlo, proponemos una soluci&#x00F3;n que es agn&#x00F3;stica a la apariencia &#x00F3;ptica del tejido y que puede ser utilizada por operadores no expertos. Utilizamos una red neuronal que predice el conjunto completo de par&#x00E1;metros mec&#x00E1;nicos de cualquier tejido utilizando datos sint&#x00E9;ticos para su entrenamiento, aprovechando una pol&#x00ED;tica extensiva de aumento de datos, aprendizaje por transferencia, mecanismos de atenci&#x00F3;n y m&#x00FA;ltiples formas de agrupamiento de activaciones internas. Una contribuci&#x00F3;n significativa de este trabajo es nuestra propuesta de una m&#x00E9;trica de similitud de propiedades mec&#x00E1;nicas en tejidos, que mide las distancias entre el comportamiento mec&#x00E1;nico de los tejidos utilizando solo im&#x00E1;genes como entrada. Validamos esta m&#x00E9;trica con un estudio de usuarios, demostrando que se relaciona en gran medida con la percepci&#x00F3;n humana. Aunque nuestro m&#x00E9;todo tiene contribuciones significativas para la estimaci&#x00F3;n del comportamiento mec&#x00E1;nico de los textiles, todav&#x00ED;a est&#x00E1; limitado en al menos</p>
<p>dos formas: utiliza im&#x00E1;genes de profundidad como entrada, que son relativamente f&#x00E1;ciles de obtener pero menos accesibles que las im&#x00E1;genes RGB, y la m&#x00E9;trica propuesta de similitud no es f&#x00E1;cilmente diferenciable, lo que dificulta su potencial como funci&#x00F3;n de p&#x00E9;rdida en problemas de optimizaci&#x00F3;n.</p>
</sec>
</sec>
<sec id="c10-app1-s4">
<label>A.4</label>
<title>Resultados</title>
<p>Estas son las principales contribuciones de esta tesis:</p>
<list list-type="bullet">
<list-item><p>Un m&#x00E9;todo basado en aprendizaje profundo para la propagaci&#x00F3;n de atributos visuales de un material a muestras m&#x00E1;s grandes del mismo material, o de un material similar. Usando una red neuronal convolucional con un n&#x00FA;mero reducido de par&#x00E1;metros, entrenada mediante un <italic>dataset fotom&#x00E9;trico</italic> y una pol&#x00ED;tica extensiva de aumentado de datos, el modelo puede generalizar a nuevas condiciones de iluminaci&#x00F3;n, c&#x00E1;mara, color o distorsiones geom&#x00E9;tricas. Ense&#x00F1;amos las capacidades de nuestro m&#x00E9;todo para la transferencia de atributos visuales de varios tipos, como las normales de la superficie del material, mapas de segmentaci&#x00F3;n sem&#x00E1;ntica o ediciones art&#x00ED;sticas. (<xref ref-type="book-part" rid="c2">Cap&#x00ED;tulo 2</xref>).</p></list-item>
<list-item><p>Una extensi&#x00F3;n al m&#x00E9;todo anterior, para habilitar un sistema de captura a dos escalas, que permite digitalizar un material a altos niveles de resoluci&#x00F3;n y precisi&#x00F3;n. Para ello, extendemos el m&#x00E9;todo presentado en el <xref ref-type="book-part" rid="c2">cap&#x00ED;tulo 2</xref> para permitir la transferencia de varios mapas de propiedades simult&#x00E1;neamente, con un mayor nivel de exactitud. (<xref ref-type="sec" rid="c2-sC">Cap&#x00ED;tulo 2.C</xref>).</p></list-item>
<list-item><p>Un modelo generativo capaz de generar <italic>texturas teselables</italic> usando una sola imagen como entrada. Explotando las capacidades de las redes generativas antag&#x00F3;nicas para la s&#x00ED;ntesis de texturas, un innovador algoritmo de manipulaci&#x00F3;n de espacios latentes, y usando el discriminador como una m&#x00E9;trica de calidad, nuestro modelo es capaz de generar texturas teselables a un mayor nivel de calidad perceptual y a un coste computacional que m&#x00E9;todos anteriores. Adem&#x00E1;s, proponemos una extensi&#x00F3;n a este m&#x00E9;todo para la <target target-type="page" id="pges_281"/>generaci&#x00F3;n de SVBRDFs teselables usando un s&#x00F3;lo material como entrada, y ense&#x00F1;amos las capacidades de nuestro algoritmo para sintetizar texturas de diferentes niveles de regularidad. (<xref ref-type="book-part" rid="c3">Cap&#x00ED;tulo 3</xref>).</p></list-item>
<list-item><p>Una representaci&#x00F3;n basada en campos neuronales capaz de codificar y transferir BTFs. Nuestro modelo se construye sobre trabajo anterior en representaciones neuronales de materiales, y proponemos una representaci&#x00F3;n latente ligera que puede ser descodificada por una red neuronal impl&#x00ED;cita en valores de reflectancia por t&#x00E9;xel. Esta representaci&#x00F3;n latente es estimada por una red neuronal del tipo <italic>autoencoder</italic>, que puede ser utilizada para propagar valores de BTF a estructuras nuevas, habilitando nuevas capacidades de edici&#x00F3;n y transferencia de materiales neuronales, proporcionando nuevas competencias a este tipo de representaciones. (<xref ref-type="book-part" rid="c4">Cap&#x00ED;tulo 4</xref>).</p></list-item>
<list-item><p>Un innovador sistema de digitalizaci&#x00F3;n de materiales capaz de, a partir de una sola imagen de entrada, generar mapas de SVBRDF a una alta resoluci&#x00F3;n, usando un esc&#x00E1;ner como dispositivo de captura. Nuestro sistema usa un conjunto de datos generado para este prop&#x00F3;sito y un novedoso modelo de translaci&#x00F3;n de imagen a imagen con mecanismos de atenci&#x00F3;n entrenado con funciones de coste espec&#x00ED;ficamente dise&#x00F1;adas para este problema. Este modelo es capaz de generar materiales digitales con un alto nivel de exactitud y detalle, usando la microgeometr&#x00ED;a del material como la principal se&#x00F1;al de supervisi&#x00F3;n. Adem&#x00E1;s, proponemos el primer algoritmo para la cuantificaci&#x00F3;n de incertidumbre en problemas de digitalizaci&#x00F3;n de materiales, construyendo sobre m&#x00E9;todos de aprendizaje profundo Bayesiano, <italic>dropout</italic> de Monte Carlo y m&#x00E9;tricas perceptuales de BRDF. Demostramos que nuestra m&#x00E9;trica es eficaz para la predicci&#x00F3;n de error en digitalizaci&#x00F3;n en el momento de la evaluaci&#x00F3;n, y para la generaci&#x00F3;n eficiente de datasets mediante aprendizaje activo. (<xref ref-type="book-part" rid="c5">Cap&#x00ED;tulo 5</xref>).</p></list-item>
<list-item><p>Un novedoso sistema de digitalizaci&#x00F3;n de propiedades mec&#x00E1;nicas de tejidos usando un sistema casual de captura e im&#x00E1;genes de profundidad. Entrenamos nuestro modelo usando &#x00FA;nicamente datos sint&#x00E9;ticos, y lo dise&#x00F1;amos para hacer uso de mecanismos de atenci&#x00F3;n neuronales, aprendizaje por transferencia, y diferentes mecanismos de agregaci&#x00F3;n de caracter&#x00ED;sticas latentes, adem&#x00E1;s de una extensiva pol&#x00ED;tica de aumentado de datos. Nuestro modelo es capaz de predecir con exactitud las propiedades mec&#x00E1;nicas de fragmentos de materiales textiles sin la necesidad de costosos sistemas de captura. Adem&#x00E1;s, proporcionamos una innovadora m&#x00E9;trica perceptual para medir la similitud entre materiales textiles en funci&#x00F3;n de sus par&#x00E1;metros mec&#x00E1;nicos, la cual validamos mediante un estudio de usuario. (<xref ref-type="book-part" rid="c6">Cap&#x00ED;tulo 6</xref>).</p></list-item>
<list-item><p>Un m&#x00E9;todo neuronal para mapas de entorno, que habilita una representaci&#x00F3;n eficiente para iluminaci&#x00F3;n global en escenas virtuales. Para ello, usamos dos redes neuronales diferentes: Un <italic>normalizing flow</italic>, capaz de aprender la distribuci&#x00F3;n de probabilidad de un mapa de entorno de entrada, y de muestrear direcciones de luz de forma eficiente; y una red neuronal impl&#x00ED;cita con activaciones sinusoidales, capaz de transformar direcciones de luz a valores de radiancia lineal en RGB. Demostramos la eficacia de nuestro <target target-type="page" id="pges_282"/>m&#x00E9;todo en aplicaciones de muestreo por importancia para renderizaci&#x00F3;n por Monte Carlo, consiguiendo representaciones de iluminaci&#x00F3;n m&#x00E1;s precisas y exactas que m&#x00E9;todos anteriores de compresi&#x00F3;n de iluminaci&#x00F3;n; y dos &#x00F3;rdenes de magnitud menor coste computacional que m&#x00E9;todos tradicionales de muestreo de iluminaci&#x00F3;n. (<xref ref-type="book-part" rid="c7">Cap&#x00ED;tulo 7</xref>).</p></list-item>
<list-item><p>Un estudio integral de las capacidades y limitaciones de las redes neuronales profundas para el problema de la decsomposici&#x00F3;n intr&#x00ED;nseca. Mediante un estudio exhaustivo del estado del arte, los datasets, funciones de coste, arquitecturas neuronales y heur&#x00ED;sticas, proporcionamos una nueva categorizaci&#x00F3;n de estos trabajos. Adem&#x00E1;s, proponemos nuevas direcciones de investigaci&#x00F3;n, construyendo sobre art&#x00ED;culos recientes en renderizaci&#x00F3;n neuronal, inverso o diferenciable; adem&#x00E1;s de los nuevos avances en sistemas de aprendizaje autom&#x00E1;tico. Esta contribuci&#x00F3;n est&#x00E1; publicada en [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>], y usamos parte de ella para la secci&#x00F3;n 1.1.</p></list-item>
</list>
<p>Estas contribuciones han dado lugar a las siguientes publicaciones:</p>
<list list-type="bullet">
<list-item><p><italic>&#x201C;Neural Photometry-guided Visual Attribute Transfer&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Elena Garces; <italic>IEEE Transactions on Visualization and Computer Graphics (TVCG), 2021</italic> [<xref ref-type="bibr" rid="CIT338">RG21</xref>].</p>
<p>Este journal tiene un factor de impacto de 5.226, su posici&#x00F3;n en el &#x00ED;ndice JCR es 13 de 110 (Q1) en la categor&#x00ED;a Computer Science, Software Engineering (datos de 2021).</p></list-item>
<list-item><p><italic>&#x201C;A Survey on Intrinsic Images: Delving Deep into Lambert and Beyond&#x201D;</italic></p>
<p>Elena Garces, Carlos Rodriguez-Pardo, Dan Casas, Jorge Lopez-Moreno; <italic>International Journal in Computer Vision (IJCV), 2022</italic> [<xref ref-type="bibr" rid="CIT104">Gar+22</xref>].</p>
<p>Este journal tiene un factor de impacto de 13.369, su posici&#x00F3;n en el &#x00ED;ndice JCR es 10 de145 (Q1) en la categor&#x00ED;a Computer Science, Artificial Intelligence (datos de 2021).</p></list-item>
<list-item><p><italic>&#x201C;SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Elena Garces; <italic>IEEE Transactions on Visualization and Computer Graphics (TVCG), 2022</italic> [<xref ref-type="bibr" rid="CIT339">RG22</xref>].</p>
<p>Este journal tiene un factor de impacto de 5.226, su posici&#x00F3;n en el &#x00ED;ndice JCR es 13 de 110 (Q1) en la categor&#x00ED;a Computer Science, Software Engineering (data from 2021)</p></list-item>
<list-item><p><italic>&#x201C;How Will It Drape Like? Capturing Fabric Mechanics From Depth Images&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Melania Prieto-Martin, Dan Casas, Elena Garces; <italic>Computer Graphics Forum (Proceedings of Eurographics 2023)</italic> [<xref ref-type="bibr" rid="CIT340">Rod+23b</xref>].</p>
<p>Este journal tiene un factor de impacto de 2.363, su posici&#x00F3;n en el &#x00ED;ndice JCR es 53 de 110 (Q2) en la categor&#x00ED;a Computer Science, Software Engineering (datos de 2021).</p></list-item>
<list-item><p><target target-type="page" id="pges_283"/><italic>&#x201C;UMat: Uncertainty-Aware Single Image High Resolution Material Capture&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Henar Dominguez, David Pascual, Elena Garces; <italic>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023</italic> [<xref ref-type="bibr" rid="CIT337">Rod+23a</xref>].</p>
<p>Esta conferencia es el evento de mayor impacto en ciencias de la computaci&#x00F3;n seg&#x00FA;n <italic>Google Scholar</italic>, y est&#x00E1; categorizada como CORE <italic>A++</italic> conference (datos de 2021)).</p></list-item>
<list-item><p><italic>&#x201C;Towards Material Digitization with a Dual-scale Optical System&#x201D;</italic></p>
<p>Elena Garces, Victor Arellano, Carlos Rodriguez-Pardo, David Pascual, Sergio Suja, Jorge Lopez-Moreno; <italic>ACM Transactions on Graphics (Proceedings of SIGGRAPH), 2023</italic> [<xref ref-type="bibr" rid="CIT103">Gar+23</xref>]. Este journal tiene un factor de impacto de 7.403, su posici&#x00F3;n en el &#x00ED;ndice JCR es 9 de 110 (Q1) en la categor&#x00ED;a Computer Science, Software Engineering (datos de 2021).</p></list-item>
<list-item><p><italic>&#x201C;NeuBTF: Neural Fields for BTF Encoding and Transfer&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Konstantinos Kazatzis, Jorge Lopez-Moreno, Elena Garces; 2023</p>
<p><italic>Esta publicaci&#x00F3;n se encuentra actualmente en revisi&#x00F3;n.</italic></p></list-item>
<list-item><p><italic>&#x201C;NEnv: Neural Environment Maps for Global Illumination&#x201D;</italic></p>
<p>Carlos Rodriguez-Pardo, Javier Fabre, Elena Garces, Jorge Lopez-Moreno; 2023 <italic>Esta publicaci&#x00F3;n se encuentra actualmente en revisi&#x00F3;n.</italic></p></list-item>
</list>
</sec>
<sec id="c10-app1-s5">
<label>A.5</label>
<title>Conclusiones</title>
<p>A lo largo de esta tesis, se han abordado diferentes retos en los campos de la inform&#x00E1;tica gr&#x00E1;fica, el aprendizaje autom&#x00E1;tico y la visi&#x00F3;n artificial; desde la generaci&#x00F3;n de representaciones neuronales m&#x00E1;s eficientes para la codificaci&#x00F3;n de componentes de las escenas, hasta el dise&#x00F1;o de algoritmos de aprendizaje profundo para la digitalizaci&#x00F3;n de materiales a bajo coste y alta precisi&#x00F3;n. En concreto, se han propuesto algoritmos para la propagaci&#x00F3;n de atributos visuales (<xref ref-type="book-part" rid="c2">Cap&#x00ED;tulo 2</xref>), para la digitalizaci&#x00F3;n realista de la reflectancia de materiales (<xref ref-type="sec" rid="c2-sC">Cap&#x00ED;tulo 2.C</xref>), su edici&#x00F3;n y s&#x00ED;ntesis hacia texturas teselables (<xref ref-type="book-part" rid="c3">Cap&#x00ED;tulo 3</xref>), o su codificaci&#x00F3;n mediante campos neuronales (<xref ref-type="book-part" rid="c4">Cap&#x00ED;tulo 4</xref>). Adem&#x00E1;s, se han propuesto sistemas de captura escalables y de alta calidad mediante dispositivos de bajo coste, incluyendo esc&#x00E1;neres para la captura de reflectancia (<xref ref-type="book-part" rid="c5">Cap&#x00ED;tulo 5</xref>) o las c&#x00E1;maras de profundidad para la captura de par&#x00E1;metros mec&#x00E1;nicos de tejidos (<xref ref-type="book-part" rid="c6">Cap&#x00ED;tulo 6</xref>). Finalmente, se ha propuesto una representaci&#x00F3;n neuronal capaz de introducir iluminaci&#x00F3;n global a escenas virtuales de forma diferenciable y computacionalmente m&#x00E1;s eficiente y precisa que las soluciones disponibles en la literatura (<xref ref-type="book-part" rid="c7">Cap&#x00ED;tulo 7</xref>). Los resultados <target target-type="page" id="pges_284"/>de esta tesis forman parte de ocho publicaciones, seis de las cuales est&#x00E1;n aceptadas en revistas acad&#x00E9;micas o conferencias internacionales, mientras que las dem&#x00E1;s se encuentran en proceso de revisi&#x00F3;n.</p>
<p>En conjunto, en esta tesis se introducen avances significativos para la representaci&#x00F3;n y generaci&#x00F3;n de escenas virtuales realistas, proporcionando m&#x00E9;todos basados en aprendizaje autom&#x00E1;tico capaces de mejorar el realismo y la eficiencia de estas escenas. El impresionante ritmo de la investigaci&#x00F3;n en inteligencia artificial dificulta prever las soluciones que se desarrollar&#x00E1;n en el futuro, sin embargo, creemos que la combinaci&#x00F3;n de soluciones basadas en datos con la renderizaci&#x00F3;n basada en la f&#x00ED;sica seguir&#x00E1; definiendo el campo en los a&#x00F1;os venideros. Esperamos que las ideas presentadas en esta tesis contribuyan a la generaci&#x00F3;n de soluciones digitales m&#x00E1;s accesibles y sostenibles medioambiental y socialmente.<target target-type="page" id="pges_285"/><target target-type="page" id="pges_286"/><target target-type="page" id="pges_287"/><target target-type="page" id="pges_288"/></p>
</sec>
</body>
</book-part>
</book-back>
</book>
