{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Ramith Hettiarachchi — Scholarly Posts",
  "description": "Curated research notes, scholarly explainers, and independent findings.",
  "user_comment": "This is the blog's curated full-text scholarly feed.",
  "home_page_url": "https://new.ramith.fyi/",
  "feed_url": "https://new.ramith.fyi/feed.json",
  "language": "en",
  "authors": [
    {
      "name": "Ramith Hettiarachchi",
      "url": "https://new.ramith.fyi/"
    }
  ],
  "items": [
    {
      "id": "https://new.ramith.fyi/posts/2024-02-12-an-open-science-journey/",
      "url": "https://new.ramith.fyi/posts/2024-02-12-an-open-science-journey/",
      "title": "An inspiring open science journey to remember 💙",
      "content_html": "<h2>An inspiring open science journey to remember 💙</h2><div data-title=\"An inspiring open science journey to remember 💙\" data-slug=\"2024-02-12-an-open-science-journey\" data-authors=\"Ramith Hettiarachchi\" data-date-published=\"2024-02-12\" data-language=\"en\" data-status=\"published\" data-license=\"CC-BY-4.0\" data-collections=\"scholarly\" data-doi=\"\"><p>By Ramith Hettiarachchi ⋅ Published <time datetime=\"2024-02-12\">2024-02-12</time></p></div><p><em><strong>Aya (Model &amp; Dataset) is going to be released in 6 hours from now. Meanwhile, I thought of writing my thoughts on my journey and the things I have learned while collaborating with so many people across the world who had a common set of values and goals. I’ve been part of this effort since June 2023, and time flew so fast that I forgot how I even joined this in the first place. So this is a reflection for me to look back again at this day, from sometime in the future.</strong></em></p><p>—</p><p><em>We all know that data is a cornerstone of AI advancements. But apart from English, most languages in the world have negligible representation on the Internet. <strong>Can we change this with the power of community collaboration? This was what <a href=\"https://cohere.com/research/aya\">Aya</a> was all about.</strong></em> It <em>was a massive global effort to uplift many under-resourced languages in the context of the current natural language advancement landscape.</em></p><p>Before elaborating on how I joined this project and the specifics of it, <strong>I want to think out loud about open science efforts.</strong> During the duration of this project, I saw,</p><ul><li>people from various age groups, various levels of expertise gather towards building datasets for under resourced languages</li><li>like minded individuals coming together for a common goal and inspiring more people to work on these research questions</li><li>people who are new to research leading this effort, carefully paying attention to very subtle details and writing a research paper encapsulating the effort in terms of a scientific article <a href=\"#ref-model\" role=\"doc-biblioref\">[1]</a> <a href=\"#ref-dataset\" role=\"doc-biblioref\">[2]</a>.</li></ul><p><strong>All this unfolding on Discord was so inspiring to watch.</strong> And to finally see a multilingual large language model (LLM), proficient in 101 languages – and more importantly, seeing Aya-101 <a href=\"#ref-model\" role=\"doc-biblioref\">[1]</a> proficient in my native language, Sinhala (සිංහල) too was such a joy 🎉❤️.</p><p>—</p><p><em>I asked these two questions from Aya in Sinhala: 1)</em> වෙහෙසකර අභියෝගයක් සාර්ථකව ජය ගත් පසු කුමක්ද කරන්න ඕනේ? 2) ලංකාවෙ නිදහස් දිනය කවද්ද?<br>While there is a long way to go for Sinhala LLMs I’m honestly suprised that it follows Sinhala instructions quite nicely!</p><figure><div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-03-31-at-23.54.26.png\" aria-label=\"Open full-size image: Aya-101 responding to a Sinhala prompt\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-03-31-at-23.54.26.png\" alt=\"Aya-101 responding to a Sinhala prompt\" loading=\"lazy\"></a></div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-03-31-at-23.53.47.png\" aria-label=\"Open full-size image: Another Sinhala conversation with Aya-101\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-03-31-at-23.53.47.png\" alt=\"Another Sinhala conversation with Aya-101\" loading=\"lazy\"></a></div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-03-31-at-23.58.08.png\" aria-label=\"Open full-size image: Sinhala instruction-following example from Aya-101\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-03-31-at-23.58.08.png\" alt=\"Sinhala instruction-following example from Aya-101\" loading=\"lazy\"></a></div></div><figcaption>(You can try the Aya-101 model here – <a href=\"https://huggingface.co/spaces/Tonic/Aya\">https://huggingface.co/spaces/Tonic/Aya</a>)</figcaption></figure><p>—</p><h3>The Start</h3><p>If I recall correctly, during my final year of undergrad, I joined C4AI (Cohere For AI<span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/4c22257bd63b615f03eeca5affc99d157539c427c7b50749233ae19938b7156a.svg\" alt=\"Typeset mathematical expression\" width=\"15.03\" height=\"10.02\"></span>) discord server. I saw so many wonderful initiatives by this community to help students/researchers who wanted to learn about AI and ultimately contribute to research. I didn’t follow these initiatives closely due to other time commitments. But every once in a while, I would check out the cool initiatives of this community.</p><p><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/4c22257bd63b615f03eeca5affc99d157539c427c7b50749233ae19938b7156a.svg\" alt=\"Typeset mathematical expression\" width=\"15.03\" height=\"10.02\"></span>Cohere for AI is a non-profit driven by researchers all over the world with the mission of solving complex machine-learning problems.</p><p>—</p><h4>Filling out a Google Form (May 2023)</h4><figure><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/IMG_8121-1.JPG\" aria-label=\"Open full-size image: A morning run in May 2023\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/IMG_8121-1.JPG\" alt=\"A morning run in May 2023\" loading=\"lazy\"></a><figcaption></figcaption></figure><p>On May 8th, I was getting ready for a morning run and saw this message on Discord.</p><figure><div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/IMG_8109-2.PNG\" aria-label=\"Open full-size image: Aya contributor signup form on a phone\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/IMG_8109-2.PNG\" alt=\"Aya contributor signup form on a phone\" loading=\"lazy\"></a></div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-02-13-at-00.40.28-3.png\" aria-label=\"Open full-size image: Discord invitation to contribute to Project Aya\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-02-13-at-00.40.28-3.png\" alt=\"Discord invitation to contribute to Project Aya\" loading=\"lazy\"></a></div></div><figcaption>(Suprisingly I had taken a screenshot of the Google Form (left picture) on that day, lucky for that, I can include it in this blog post!)</figcaption></figure><p>—</p><p>When filling something up, I always think whether it is something I can do given my other commitments<span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/146c579653eb92cf0b59169294d3ac66bf20b08144752d106157babc3e6fd0b9.svg\" alt=\"Typeset mathematical expression\" width=\"6.51\" height=\"10.02\"></span> 😅</p><p>I knew I didn’t have much time to dedicate, but the message above seemed reasonable so I filled it out. It seemed like only knowing the language was enough. Words like<code>multilingual</code>and<code>underrepresented languages</code>made it so tempting to fill out the Google form.</p><p><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/146c579653eb92cf0b59169294d3ac66bf20b08144752d106157babc3e6fd0b9.svg\" alt=\"Typeset mathematical expression\" width=\"6.51\" height=\"10.02\"></span> During undergrad, there are many occasions when I took too many things to my plate..</p><p>During high-school I used to make apps for Sinhala transliteration-<a href=\"https://archive.org/details/sinhala\">link</a> . So I have a softspot for tool that advance the usage of my native language.</p><p>—</p><h4>An Email ✉️</h4><p>June, 2023 was a great month, I was finishing up some experiments for my ICML 2023 workshop paper, and writing the paper. And then I received this email..</p><figure><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-02-13-at-00.06.32.png\" aria-label=\"Open full-size image: Email invitation to represent Sinhala in Project Aya\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-02-13-at-00.06.32.png\" alt=\"Email invitation to represent Sinhala in Project Aya\" loading=\"lazy\"></a><figcaption></figcaption></figure><p>it was sort of unexpected.. felt serious, but then again I felt like it’s something I’m capable of doing during my free time (a weekly meeting, spreading info about this project sounded good and doable 🤷🏻‍♂️).</p><p>So I replied that I’d want to help out to represent my native language! <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/bd6f74fe981f6c83f3ad204b96d64df2093aeeb5953241441316cdf821a192db.svg\" alt=\"Typeset mathematical expression\" width=\"6.51\" height=\"10.02\"></span>.</p><p><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/bd6f74fe981f6c83f3ad204b96d64df2093aeeb5953241441316cdf821a192db.svg\" alt=\"Typeset mathematical expression\" width=\"6.51\" height=\"10.02\"></span> After all I was missing volunteering after completing my undergrad too… so felt that this would be nice..</p><h3>Journey as a Language Ambassador</h3><p>I attended the first meeting where I got introduced to Project Aya. The presentation was incredible… I thought to myself, “a great research group with a clear vision of what they want to achieve in 2023..”. There was so much positivity and hope within this community, and I loved that energy!</p><h4>Somewhat of a rough start for Sinhala 🤕</h4><p>Now it was my turn to contribute to my language.  I remember going to the Aya annotation platform<span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/d0ca01ab81c021915a2f3e3d0d35466c50ebf756c4d7989c5539ab714c84e30c.svg\" alt=\"Typeset mathematical expression\" width=\"7.7\" height=\"10.02\"></span>, and filling out my details (such as languages I speak).</p><p>I was surprised to see Sinhala prompts and completions 😃, but soon I realized that they are machine translations which most often did not make much sense 🤕 (see left side of the below picture for some context).</p><p><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/d0ca01ab81c021915a2f3e3d0d35466c50ebf756c4d7989c5539ab714c84e30c.svg\" alt=\"Typeset mathematical expression\" width=\"7.7\" height=\"10.02\"></span> Where human contributions are collected.</p><figure><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-02-18-at-16.39.02.png\" aria-label=\"Open full-size image: Sinhala annotation examples on the Aya platform\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-02-18-at-16.39.02.png\" alt=\"Sinhala annotation examples on the Aya platform\" loading=\"lazy\"></a><figcaption></figcaption></figure><p>So, we gradually started to refine those. Soon with the help of many Sinhala contributors, we got a massive pool of prompts and completions that we kept on refining.</p><h4>Realizing that we need to speed up</h4><p>From June to the end of September, we gained a lot of contributions, but soon, we realized our pace was too slow 😬. So what we did was to visualize <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/d0ca01ab81c021915a2f3e3d0d35466c50ebf756c4d7989c5539ab714c84e30c.svg\" alt=\"Typeset mathematical expression\" width=\"7.7\" height=\"10.02\"></span> our goal and work towards it. Furthermore, we spread out the message through many contributors (kudos to Jalina, Chamod, Nawoda, and Chanuka)</p><p><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/d0ca01ab81c021915a2f3e3d0d35466c50ebf756c4d7989c5539ab714c84e30c.svg\" alt=\"Typeset mathematical expression\" width=\"7.7\" height=\"10.02\"></span> Our goal was to reach 5000 re-annotations and 4000 original annotations</p><figure><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-02-18-at-17.25.57.png\" aria-label=\"Open full-size image: Discord message outlining Sinhala annotation goals\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Screenshot-2024-02-18-at-17.25.57.png\" alt=\"Discord message outlining Sinhala annotation goals\" loading=\"lazy\"></a><figcaption>(a discord message from the Aya Server)</figcaption></figure><h3>So we accelerated our pace! 🇱🇰</h3><p>For the next remaining 3 months, all Sinhala contributors helped immensely to reach our goals – we even surpassed our original goals!</p><p>The pictures below capture some of the initiatives we took to gather more contributors..</p><figure><div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Aya---Dec-15---Closing-the-Contribution-Chapter-30--dragged-.png\" aria-label=\"Open full-size image: Sinhala contributor outreach presentation, slide 30\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Aya---Dec-15---Closing-the-Contribution-Chapter-30--dragged-.png\" alt=\"Sinhala contributor outreach presentation, slide 30\" loading=\"lazy\"></a></div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Aya---Dec-15---Closing-the-Contribution-Chapter-31--dragged-.png\" aria-label=\"Open full-size image: Sinhala contributor outreach presentation, slide 31\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Aya---Dec-15---Closing-the-Contribution-Chapter-31--dragged-.png\" alt=\"Sinhala contributor outreach presentation, slide 31\" loading=\"lazy\"></a></div></div><div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Aya---Dec-15---Closing-the-Contribution-Chapter-32--dragged-.png\" aria-label=\"Open full-size image: Sinhala contributor outreach presentation, slide 32\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Aya---Dec-15---Closing-the-Contribution-Chapter-32--dragged-.png\" alt=\"Sinhala contributor outreach presentation, slide 32\" loading=\"lazy\"></a></div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Aya---Dec-15---Closing-the-Contribution-Chapter-33--dragged-.png\" aria-label=\"Open full-size image: Sinhala contributor outreach presentation, slide 33\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/Aya---Dec-15---Closing-the-Contribution-Chapter-33--dragged-.png\" alt=\"Sinhala contributor outreach presentation, slide 33\" loading=\"lazy\"></a></div></div><figcaption></figcaption></figure><p>—</p><figure><div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/IMG_5904-1.webp\" aria-label=\"Open full-size image: Commemorative chip with Aya Sinhala contributor names\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/IMG_5904-1.webp\" alt=\"Commemorative chip with Aya Sinhala contributor names\" loading=\"lazy\"></a></div><div><a href=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/IMG_5886.webp\" aria-label=\"Open full-size image: Another view of the Aya Sinhala contributors chip\"><img src=\"https://new.ramith.fyi/assets/posts/2024-02-12-an-open-science-journey/IMG_5886.webp\" alt=\"Another view of the Aya Sinhala contributors chip\" loading=\"lazy\"></a></div></div><figcaption>(a 1x1 inch chip fabricated with the names of Aya Sinhala Contributors as a memory)</figcaption></figure><section role=\"doc-bibliography\"><h2>References</h2><ol><li id=\"ref-model\" data-paper-title=\"Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model\"><a href=\"https://arxiv.org/abs/2402.07827\">“Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.”</a></li><li id=\"ref-dataset\" data-paper-title=\"Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning\"><a href=\"https://arxiv.org/abs/2402.06619\">“Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning.”</a></li></ol></section><footer>Unless otherwise noted, the original text of this post is licensed under the <a href=\"https://creativecommons.org/licenses/by/4.0/\">Creative Commons Attribution 4.0 International License</a>. Separately credited material may have different rights. <a href=\"https://new.ramith.fyi/license/\">Licensing details</a>.</footer>",
      "date_published": "2024-02-12T00:00:00Z",
      "authors": [
        {
          "name": "Ramith Hettiarachchi",
          "url": "https://new.ramith.fyi"
        }
      ],
      "language": "en",
      "tags": [
        "scholarly"
      ]
    },
    {
      "id": "https://new.ramith.fyi/posts/2023-06-24-differentiable-search-evolutionary-trees/",
      "url": "https://new.ramith.fyi/posts/2023-06-24-differentiable-search-evolutionary-trees/",
      "title": "Differentiable Search of Evolutionary Trees from Leaves — ICML 2023 Workshop Paper",
      "content_html": "<h2>Differentiable Search of Evolutionary Trees from Leaves — ICML 2023 Workshop Paper</h2><div data-title=\"Differentiable Search of Evolutionary Trees from Leaves — ICML 2023 Workshop Paper\" data-slug=\"2023-06-24-differentiable-search-evolutionary-trees\" data-authors=\"Ramith Hettiarachchi\" data-date-published=\"2023-06-24\" data-language=\"en\" data-status=\"published\" data-license=\"CC-BY-4.0\" data-collections=\"scholarly; workshop paper; phylogenetics; differentiable programming\" data-doi=\"\"><p>By Ramith Hettiarachchi ⋅ Published <time datetime=\"2023-06-24\">2023-06-24</time></p></div><blockquote aria-label=\"ICML 2023 workshop paper and resources\"><p>ICML 2023 workshop paper</p><h3>Differentiable Search of Evolutionary Trees from Leaves</h3><p>Ramith Hettiarachchi, Avi Z. Swartz, Sergey Ovchinnikov</p><p>ICML 2023 workshops: Differentiable Almost Everything (DiffAE) and Sampling and Optimization in Discrete Space (SODS)</p><div>This page collects the workshop paper and its supporting resources. The paper contains the complete method, experiments, and claims.</div><p aria-label=\"Research artifacts\"><a href=\"https://doi.org/10.1101/2023.07.23.550206\">Read paper ↗</a> <a href=\"https://ramith.fyi/assets/pdf/Diff-Evol-Trees_ICML.pdf\">View poster ↗</a> <a href=\"https://github.com/ramithuh/diff-evol-tree-search\">Browse code ↗</a> <a href=\"https://ericmjl.github.io/blog/2023/8/7/journal-club-differentiable-search-of-evolutionary-trees/\">External explainer ↗</a> </p></blockquote><blockquote><div aria-hidden=\"true\">💡</div><div><strong>In one sentence:</strong> We relax the discrete choices behind a phylogenetic tree so that its topology and unknown ancestral sequences can be improved together using gradients.</div></blockquote><h3>Related paper</h3><p>Hettiarachchi, R., Swartz, A. Z., and Ovchinnikov, S. “Differentiable Search of Evolutionary Trees from Leaves.” ICML 2023 Workshops: DiffAE and SODS. DOI: <a href=\"https://doi.org/10.1101/2023.07.23.550206\">10.1101/2023.07.23.550206</a></p><footer>Unless otherwise noted, the original text of this post is licensed under the <a href=\"https://creativecommons.org/licenses/by/4.0/\">Creative Commons Attribution 4.0 International License</a>. Separately credited material may have different rights. <a href=\"https://new.ramith.fyi/license/\">Licensing details</a>.</footer>",
      "date_published": "2023-06-24T00:00:00Z",
      "authors": [
        {
          "name": "Ramith Hettiarachchi",
          "url": "https://new.ramith.fyi"
        }
      ],
      "language": "en",
      "tags": [
        "scholarly",
        "workshop paper",
        "phylogenetics",
        "differentiable programming"
      ]
    },
    {
      "id": "https://new.ramith.fyi/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/",
      "url": "https://new.ramith.fyi/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/",
      "title": "ESM-2 (evolutionary-scale prediction of atomic level protein structure with a language model)",
      "content_html": "<h2>ESM-2 (evolutionary-scale prediction of atomic level protein structure with a language model)</h2><div data-title=\"ESM-2 (evolutionary-scale prediction of atomic level protein structure with a language model)\" data-slug=\"2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model\" data-authors=\"Ramith Hettiarachchi\" data-date-published=\"2023-02-10\" data-language=\"en\" data-status=\"published\" data-license=\"CC-BY-4.0\" data-collections=\"scholarly\" data-doi=\"\"><p>By Ramith Hettiarachchi ⋅ Published <time datetime=\"2023-02-10\">2023-02-10</time></p></div><blockquote><div aria-hidden=\"true\">✏️</div><div><em>I write these paper “summaries” for me to clearly understand the paper by summarizing the paper and synthesizing the literature. I hope they might be helpful to some readers too. If you have feedback, please write to <a href=\"mailto:hello@ramith.fyi\">hello@ramith.fyi</a>.</em></div></blockquote><h3>Highlights of the ESM-2 Paper</h3><blockquote><div aria-hidden=\"true\">💡</div><div><ul><li>Train protein language models up to 15B parameters.<sup><a href=\"#fn-1\" id=\"fnref-1\">[Note 1]</a></sup></li><li>Infer structure directly from a primary sequence using a language model.</li><li>Leverage evolutionary patterns captured by the language model to produce atomic-level predictions.</li><li>Achieve an order-of-magnitude speedup—up to 60×—in high-resolution structure prediction.</li><li>Present the ESM Metagenomic Atlas: structural characterization of more than 617 million metagenomic proteins.<sup><a href=\"#fn-2\" id=\"fnref-2\">[Note 2]</a></sup><sup><a href=\"#fn-3\" id=\"fnref-3\">[Note 3]</a></sup></li></ul></div></blockquote><h3>1. Introduction</h3><h4>1.1 — Structure and function are hidden in sequences</h4><p>Biological properties of proteins influence which positions in a sequence can undergo mutations. From these observations, we can identify evolutionary patterns such as coevolution and conservation of amino acids. These patterns can help us infer properties of protein function and structure.<sup><a href=\"#fn-4\" id=\"fnref-4\">[Note 4]</a></sup></p><p>Usually, we align sequences before drawing conclusions about function and structure. This intermediate representation, known as a <em>multiple sequence alignment (MSA)</em>, has a high time complexity because we must first search for related sequences and then align them.<sup><a href=\"#fn-5\" id=\"fnref-5\">[Note 5]</a></sup></p><p>What if we can get rid of this intermediate representation? That’s one aspect this paper accomplishes.</p><h4>1.2 — Large language models (LLMs)</h4><p>Historically, language models were pretrained using objectives such as predicting the next word in a sentence. Devlin et al.‘s <a href=\"https://arxiv.org/abs/1810.04805\">BERT</a> showed that masking some input tokens and predicting them—the masked-language-model objective, or MLM—is an effective pretraining strategy.<sup><a href=\"#fn-6\" id=\"fnref-6\">[Note 6]</a></sup></p><h4>1.3 — Contributions</h4><p>Inspired by this widely adopted strategy, the authors hypothesise that filling in missing amino acids might produce representations rich enough to infer structure. They therefore <strong>scale protein language models from 8 million parameters up to 15 billion parameters</strong>. Doing so reveals the following:</p><ul><li>Atomic-level structure prediction directly from sequence.</li><li>A strong correlation between perplexity and structure-prediction accuracy.</li><li>Up to 60× faster inference.</li><li><strong>No need to search</strong> for related sequences.</li></ul><p>Because this approach improves speed by one to two orders of magnitude and does not need an MSA, the authors expand structure prediction to the much larger and more diverse space of metagenomic proteins. In summary, they:</p><ul><li><p>Predict structures for more than 617 million sequences in MGnify90.<sup><a href=\"#fn-7\" id=\"fnref-7\">[Note 7]</a></sup></p><ul><li><p>Of those structures, 225 million are high-confidence predictions.</p><ul><li>76.8% are disjoint from UniRef90 at 90% sequence identity.</li><li>12.6% have no experimental ground truth.</li></ul></li></ul></li></ul><div><p>Open <a href=\"https://esmatlas.com/explore\">ESM Metagenomic Atlas explorer</a> in a new tab.</p></div><h3>2. Method</h3><h4>2.1 — How does structure emerge from a language model trained on sequences?</h4><p>The ESM-2 language model is trained with approximately 65 million unique sequences.<sup><a href=\"#fn-8\" id=\"fnref-8\">[Note 8]</a></sup> Because the MLM objective asks the model to predict missing amino acids using their neighboring context, the model must learn interdependencies between amino acids. Previous work—<a href=\"#ref-transformer\" role=\"doc-biblioref\">[2]</a> and <a href=\"#ref-contacts\" role=\"doc-biblioref\">[3]</a>—showed that Transformer models trained with MLM on protein sequences develop attention patterns that correspond to residue-residue contact maps.<sup><a href=\"#fn-9\" id=\"fnref-9\">[Note 9]</a></sup></p><p>After training the language model, the authors use the approach from <a href=\"#ref-contacts\" role=\"doc-biblioref\">[3]</a> to compute contact maps from attention patterns. A logistic regression identifies contacts as follows.</p><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/contact-prediction-pipeline.png\" aria-label=\"Open full-size image: Annotated contact-prediction pipeline from Transformer attention maps\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/contact-prediction-pipeline.png\" alt=\"Annotated contact-prediction pipeline from Transformer attention maps\" loading=\"lazy\"></a><figcaption>Attention maps from a masked protein language model are symmetrized and passed to logistic regression to predict residue-residue contacts. Source figure: <a href=\"https://doi.org/10.1101/2020.12.15.422761\">Rao et al. (2020)</a> annotations added for this article.</figcaption></figure><h4>2.2 — What about atomic-level structure? (Enter ESMFold)</h4><p>The authors extract contact maps from attention patterns, but predicting atomic spatial coordinates requires an equivariant Transformer. They use the <a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/alphafold-structure-module.png\" data-image-alt=\"Annotated AlphaFold structure module and invariant point attention diagram\" data-image-caption=\"AlphaFold structure module and invariant point attention.\"><em>structure module</em></a> introduced in <a href=\"https://doi.org/10.1038/s41586-021-03819-2\">AlphaFold</a>. This module projects atomic spatial coordinates from the language model’s internal representation. The complete architecture is called ESMFold.</p><p><em>Steps in ESMFold</em></p><ol><li value=\"1\">Process the sequence through ESM-2.</li><li value=\"2\">Pass the representation learned by ESM-2 through a series of <em>folding blocks</em>. Each block <a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-folding-trunk-code.png\" data-image-alt=\"Annotated ESMFold folding-trunk and triangular-attention source code\" data-image-caption=\"Folding-trunk code mapped to the triangular self-attention block.\">sequentially updates</a> a sequence representation and a pairwise representation.</li><li value=\"3\">Pass the result to the structure module.</li><li value=\"4\">Repeat with three recycling steps. <a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-recycling-code.png\" data-image-alt=\"Annotated ESMFold trunk code showing folding, structure prediction, and recycling\" data-image-caption=\"ESMFold folding-block iteration, structure-module call, and recycling.\">View annotated code.</a></li></ol><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-architecture.png\" aria-label=\"Open full-size image: ESMFold architecture with ESM-2, folding trunk, structure module, and recycling\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-architecture.png\" alt=\"ESMFold architecture with ESM-2, folding trunk, structure module, and recycling\" loading=\"lazy\"></a><figcaption>ESMFold architecture: ESM-2 representations pass through a folding trunk and an AlphaFold-derived structure module, with recycling. Source figure: <a href=\"https://doi.org/10.1101/2022.07.20.500902\">Lin et al. (2022)</a> annotations added for this article.</figcaption></figure><p>The supporting figures connect the structure-module diagram to implementation details. Select any image to open the full-resolution version.</p><div><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/alphafold-structure-module.png\" aria-label=\"Open full-size image: Annotated AlphaFold structure module and invariant point attention diagram\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/alphafold-structure-module.png\" alt=\"Annotated AlphaFold structure module and invariant point attention diagram\" loading=\"lazy\"></a><figcaption>AlphaFold’s structure module and invariant point attention. Source figures: <a href=\"https://doi.org/10.1038/s41586-021-03819-2\">Jumper et al. (2021)</a>.</figcaption></figure><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-folding-trunk-code.png\" aria-label=\"Open full-size image: Annotated ESMFold folding-trunk and triangular-attention source code\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-folding-trunk-code.png\" alt=\"Annotated ESMFold folding-trunk and triangular-attention source code\" loading=\"lazy\"></a><figcaption>Mapping the folding trunk to its triangular self-attention block.</figcaption></figure><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-recycling-code.png\" aria-label=\"Open full-size image: Annotated ESMFold trunk code showing folding, structure prediction, and recycling\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-recycling-code.png\" alt=\"Annotated ESMFold trunk code showing folding, structure prediction, and recycling\" loading=\"lazy\"></a><figcaption>Folding-block iteration, the structure-module call, and recycling.</figcaption></figure></div><p><strong>Training:</strong> To train the structure model to obtain spatial coordinates, the authors use experimentally determined structures from the <a href=\"https://www.rcsb.org/\">Protein Data Bank</a>—approximately 25,000 clusters covering around 325,000 structures. This is augmented with 12 million structures predicted by AlphaFold2.<sup><a href=\"#fn-10\" id=\"fnref-10\">[Note 10]</a></sup></p><p><strong>Evaluation:</strong> 194 CAMEO proteins and 51 CASP14 proteins.</p><p>This language model based approach vastly simplifies the usual SOTA structure prediction process by eliminating the need for the following <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/bd6f74fe981f6c83f3ad204b96d64df2093aeeb5953241441316cdf821a192db.svg\" alt=\"Typeset mathematical expression\" width=\"6.51\" height=\"10.02\"></span>,</p><ul><li>External evolutionary databases</li><li>Multiple sequence alignments (MSAs)</li><li>Templates</li></ul><p>For example, <a href=\"https://doi.org/10.1038/s41586-021-03819-2\">AlphaFold</a> requires access to these resources.</p><h3>3. Results</h3><h4>3.1 — How well does it predict structures?</h4><p>As mentioned before, they evaluate performance on CAMEO and CASP14 proteins and check how well the structure was predicted using the TM-Score.</p><p>In predicting the structure just by single sequences, ESMFold achieves very good performance compared to AlphaFold and RoseTTAFold.</p><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-structure-performance.png\" aria-label=\"Open full-size image: ESMFold, AlphaFold2, and RoseTTAFold TM-score comparisons on CAMEO and CASP14\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-structure-performance.png\" alt=\"ESMFold, AlphaFold2, and RoseTTAFold TM-score comparisons on CAMEO and CASP14\" loading=\"lazy\"></a><figcaption>Single-sequence and full-pipeline TM-score comparisons on CAMEO and CASP14; point color encodes perplexity. Source: Figure 2B in <a href=\"https://doi.org/10.1101/2022.07.20.500902\">Lin et al. (2022)</a> annotations added for this article.</figcaption></figure><h4>3.2 — How important is the language model in the pipeline?</h4><p>The key question that arises is <strong>how important is the representation learnt by the LM for the task of structure prediction</strong> . To quantify this we need several metrics.</p><p>First, we need to characterize how well the language model understands protein sequences. This is where perplexity comes in. We already have the TM-score to measure how closely a predicted structure matches the ground truth.</p><p><strong>Thus, the graph to the right in Fig. 2B shows that,</strong></p><ul><li><strong>High ESMFold TM-Scores</strong> have <strong>low perplexity</strong> scores (numerically speaking, on CAMEO, Pearson correlation coefficient is −0.55 and in CASP14 it’s −0.67)</li></ul><blockquote><div aria-hidden=\"true\">✏️</div><div><p><strong>What’s perplexity?</strong></p><p>Perplexity describes how well a language model can predict a protein sequence. It ranges from 1 for a perfect model to 20 for random predictions; intuitively, it is the number of amino acids the model is choosing between at each prediction.</p></div></blockquote><h5>How can we achieve lower perplexity?</h5><p>We now know that a <em>better language model representation</em>—with lower perplexity—leads to better structure prediction. How can we obtain a better representation? 🤔 Is scaling all you need?</p><p>To answer this question, authors explore the effect of scaling and look at what happens to the following :</p><ul><li>Precision @ L</li><li>Change in perplexity</li></ul><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esm2-scaling-contact-precision.png\" aria-label=\"Open full-size image: Scatter plots comparing long-range contact precision across ESM-2 model sizes\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esm2-scaling-contact-precision.png\" alt=\"Scatter plots comparing long-range contact precision across ESM-2 model sizes\" loading=\"lazy\"></a><figcaption>Effect of scaling ESM-2 on long-range contact precision; point color shows the change in perplexity. Source: Figure 1D in <a href=\"https://doi.org/10.1101/2022.07.20.500902\">Lin et al. (2022)</a> annotations added for this article.</figcaption></figure><p>The authors plot how long-range precision at L changes when moving from a smaller model on the x-axis to a larger model on the y-axis. Points above the diagonal suggest that scaling improves <em>long-range precision at L</em> for some proteins.</p><h5>Is scaling the answer?</h5><p>It is not that simple. Although P@L increases with scale for some proteins, the number of evolutionarily related sequences tells another story. Language models struggle when less relevant training data is available for a query. This is intuitive—more study, better results—but does it reflect memorization rather than understanding?</p><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esm2-sequence-depth-contact-precision.png\" aria-label=\"Open full-size image: Long-range contact precision plotted against the number of related sequences\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esm2-sequence-depth-contact-precision.png\" alt=\"Long-range contact precision plotted against the number of related sequences\" loading=\"lazy\"></a><figcaption>Long-range precision at L as a function of the number of related sequences for several ESM-2 model sizes. Source: <a href=\"https://doi.org/10.1101/2022.07.20.500902\">Lin et al. (2022)</a> annotations added for this article.</figcaption></figure><p>The authors could have used a different color scheme: white points indicate no change in perplexity, while a point on the diagonal indicates no improvement in long-range precision at L.<sup><a href=\"#fn-11\" id=\"fnref-11\">[Note 11]</a></sup></p><h4>3.3 — What about prediction speed?</h4><ul><li>A protein with 384 residues on one NVIDIA V100 GPU: 14.2 seconds.<sup><a href=\"#fn-12\" id=\"fnref-12\">[Note 12]</a></sup></li><li>Shorter sequences  60x speedup</li></ul><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-inference-speed.png\" aria-label=\"Open full-size image: Log-scale inference time comparison for ESMFold, AlphaFold, and RoseTTAFold\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esmfold-inference-speed.png\" alt=\"Log-scale inference time comparison for ESMFold, AlphaFold, and RoseTTAFold\" loading=\"lazy\"></a><figcaption>Inference time versus sequence length. ESMFold is fastest for shorter sequences, while its pair representation makes long sequences more expensive. MSA search time for the other methods is not included.</figcaption></figure><h4>3.4 — Comparison with other protein language models</h4><figure><a href=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esm2-model-comparison.png\" aria-label=\"Open full-size image: Table comparing ESM-2 sizes and other protein language models\"><img src=\"https://new.ramith.fyi/assets/posts/2023-02-10-esm-2-evolutionary-scale-prediction-of-atomic-level-protein-structure-with-a-language-model/esm2-model-comparison.png\" alt=\"Table comparing ESM-2 sizes and other protein language models\" loading=\"lazy\"></a><figcaption>Validation perplexity, long-range contact precision, and structure-prediction results across ESM-2 sizes and comparison protein language models. Source: <a href=\"https://doi.org/10.1101/2022.07.20.500902\">Lin et al. (2022)</a> annotations added for this article.</figcaption></figure><h3>4. Conclusion</h3><p>It’s remarkable that these authors scale protein language models and it has resulted in learning structure hidden through databases of sequences, and thus we do not need to depend onto the MSA.</p><p>Is it because, the model has learnt to obtain the signal which we previously obtained through MSAs? What can we tell about the performance of sequences that had less number of evolutionary sequences in training data? why does it still struggle to obtain decent performance. It would be very interesting to analyze these directions.</p><p>Thanks for reading this, hope you found it useful. If you have any suggestions/ comments please share below.</p><section role=\"doc-bibliography\"><h2>References</h2><ol><li id=\"ref-esm\" data-paper-title=\"Evolutionary-scale prediction of atomic level protein structure with a language model\"><a href=\"https://doi.org/10.1101/2022.07.20.500902\">“Evolutionary-scale prediction of atomic level protein structure with a language model.”</a></li><li id=\"ref-transformer\" data-paper-title=\"BERTology Meets Biology: Interpreting Attention in Protein Language Models\"><a href=\"https://arxiv.org/abs/2006.15222\">“BERTology Meets Biology: Interpreting Attention in Protein Language Models.”</a></li><li id=\"ref-contacts\" data-paper-title=\"Transformer protein language models are unsupervised structure learners\"><a href=\"https://doi.org/10.1101/2020.12.15.422761\">“Transformer protein language models are unsupervised structure learners.”</a></li><li id=\"ref-discussion\" data-paper-title=\"Eric Wallace: discussion of protein language models\"><a href=\"https://twitter.com/Eric_Wallace_/status/1592929060539469824\">“Eric Wallace: discussion of protein language models.”</a></li></ol></section><footer>Unless otherwise noted, the original text of this post is licensed under the <a href=\"https://creativecommons.org/licenses/by/4.0/\">Creative Commons Attribution 4.0 International License</a>. Separately credited material may have different rights. <a href=\"https://new.ramith.fyi/license/\">Licensing details</a>.</footer><h2>Notes</h2><ol><li id=\"fn-1\"><sup><a href=\"#fnref-1\">1</a></sup> The largest language model for proteins at the time.</li><li id=\"fn-2\"><sup><a href=\"#fnref-2\">2</a></sup> Of these, 225 million were high-confidence predictions. See the <a href=\"https://esmatlas.com/\">ESM Metagenomic Atlas</a>.</li><li id=\"fn-3\"><sup><a href=\"#fnref-3\">3</a></sup> Metagenomics describes both sequencing DNA purified directly from a natural environment and the research field studying microbial communities in their natural state (Godzik, 2011).</li><li id=\"fn-4\"><sup><a href=\"#fnref-4\">4</a></sup> In essence, information about protein structure and function is hidden in sequences.</li><li id=\"fn-5\"><sup><a href=\"#fnref-5\">5</a></sup> The search process in AlphaFold and RoseTTAFold pipelines can take more than ten minutes.</li><li id=\"fn-6\"><sup><a href=\"#fnref-6\">6</a></sup> Unlike left-to-right pretraining, the MLM objective lets the representation combine both left and right context, enabling a deeply bidirectional Transformer (paraphrased from the BERT paper).</li><li id=\"fn-7\"><sup><a href=\"#fnref-7\">7</a></sup> This took two weeks on 2,000 GPUs. 🤯</li><li id=\"fn-8\"><sup><a href=\"#fnref-8\">8</a></sup> The sequences were sampled with even weighting across approximately 43 million UniRef50 training clusters.</li><li id=\"fn-9\"><sup><a href=\"#fnref-9\">9</a></sup> I plan to write a separate article about this work.</li><li id=\"fn-10\"><sup><a href=\"#fnref-10\">10</a></sup> During training, predicted structures are sampled 75% of the time and real structures 25% of the time.</li><li id=\"fn-11\"><sup><a href=\"#fnref-11\">11</a></sup> It would be useful if the plot made the proportion of proteins with improved performance directly visible, although that might make it too cluttered.</li><li id=\"fn-12\"><sup><a href=\"#fnref-12\">12</a></sup> This is a 6× speedup compared with AlphaFold2.</li></ol>",
      "date_published": "2023-02-10T00:00:00Z",
      "authors": [
        {
          "name": "Ramith Hettiarachchi",
          "url": "https://new.ramith.fyi"
        }
      ],
      "language": "en",
      "tags": [
        "scholarly"
      ]
    },
    {
      "id": "https://new.ramith.fyi/posts/2022-01-26-deep-residual-learning-for-image-recognition/",
      "url": "https://new.ramith.fyi/posts/2022-01-26-deep-residual-learning-for-image-recognition/",
      "title": "Deep Residual Learning for Image Recognition",
      "content_html": "<h2>Deep Residual Learning for Image Recognition</h2><div data-title=\"Deep Residual Learning for Image Recognition\" data-slug=\"2022-01-26-deep-residual-learning-for-image-recognition\" data-authors=\"Ramith Hettiarachchi\" data-date-published=\"2022-01-26\" data-language=\"en\" data-status=\"published\" data-license=\"CC-BY-4.0\" data-collections=\"scholarly\" data-doi=\"\"><p>By Ramith Hettiarachchi ⋅ Published <time datetime=\"2022-01-26\">2022-01-26</time></p></div><h3>Highlights of ResNet Paper</h3><p>Summary of He et al. <a href=\"#ref-resnet\" role=\"doc-biblioref\">[1]</a>.</p><ul><li>Present a <strong>residual learning framework</strong> to train very deep networks more easily.</li><li>Training 8x deeper networks than VGG-net</li><li>3.57% error on the ImageNet test set, 28% relative improvement on COCO object detection dataset.</li><li>Topping the leader board in ILSVRC &amp; COCO 2015 (ImageNet classification, detection, localization, COCO detection &amp; segmentation)</li></ul><h3>Introduction</h3><p>From the ImageNet Classification results in 2014-2015, it was evident that having deeper networks helps to learn greater levels of features. As shown in the figure below, we can see that VGG-Net and GoogLeNet have reduced the top-5 error rate further by having deeper networks.</p><p>So if we just stack more and more layers, does that help? Turns out it doesn’t. One reason for this is the <a href=\"https://people.idsia.ch/~juergen/fundamentaldeeplearningproblem.html\">vanishing gradient problem</a> which was studied by Sepp Hochreiter in 1991 and discussed over the years <a href=\"#ref-gradient\" role=\"doc-biblioref\">[2]</a>. This problem makes it difficult for a network to converge from the start. This issue has been addressed through various initialization methods and through batch normalization <a href=\"#ref-batchnorm\" role=\"doc-biblioref\">[3]</a>.</p><p>Even when a network starts converging, in deeper networks, researchers have found that there is a degradation of accuracy. Particularly, <em>when we start increasing the depth</em> of a model the accuracy gets saturated, and then <strong><em>it degrades rapidly</em></strong> . As  evident from the learning curves, this is not due to overfitting. Check the training-error curves in the left figure as depth increases in increments of 12.</p><p>—</p><figure><p><sup><a href=\"#feed-note-1\">[Note 1]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/f085d0ad88fba5b5ea9420806762e1bcb78d3e0f3d6a9539ee383c1d830bc98e.png\"></figure><figure><p><sup><a href=\"#feed-note-2\">[Note 2]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/8a6f245766e7a4c1fad9e3491fcd5e7218ab167491305f797475df00f24d3c4d.png\"></figure><p>—</p><p>The authors of the ResNet paper argue that, even if we increase the depth, theoretically there should be solution which gives the same accuracy. So it’s basically the shallow network + layers with identity transform.💡 Why can’t the added layers become an <strong>identity mapping 🤷🏻</strong> However, the problem seems that the optimizers cannot reach that solution. So can we do a trick and get there easily? That’s what authors hypothesize.</p><h3>Methodology - Deep Residual Learning</h3><h4>Fitting a residual mapping</h4><p><strong><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/83b116716e47411f1f1fbd5f322c8c4ffa04964c4591199d09ce96a740566535.svg\" alt=\"Typeset mathematical expression\" width=\"12.53\" height=\"10.02\"></span></strong> - Mapping that needs to be fit by <strong>few</strong> stacked layers<br><br><strong><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/f03df5b164ae6118ebf27ba0c88ac57a952ee2d29e1540de3d37a6da6067b465.svg\" alt=\"Typeset mathematical expression\" width=\"7.98\" height=\"10.02\"></span></strong> - input to the first of those layers</p><p><em>Let’s say we need to approximate the function</em> <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/83b116716e47411f1f1fbd5f322c8c4ffa04964c4591199d09ce96a740566535.svg\" alt=\"Typeset mathematical expression\" width=\"12.53\" height=\"10.02\"></span> by some set of layers of a neural network. The authors propose that, rather than learning <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/83b116716e47411f1f1fbd5f322c8c4ffa04964c4591199d09ce96a740566535.svg\" alt=\"Typeset mathematical expression\" width=\"12.53\" height=\"10.02\"></span>, the <em>few layers</em> approximate a <em>residual function</em> (<span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/e9f9cb426972b56e08bc6efba41895ebfa06dae9e281c00786ce2352e5b5a57c.svg\" alt=\"Typeset mathematical expression\" width=\"38.43\" height=\"10.02\"></span>). 🤔</p><p>Let’s denote this new function by <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/66e75a1dedfe2609c4ae1067eb296cb00e8e9515b2533a5181a3bba32d4b22a5.svg\" alt=\"Typeset mathematical expression\" width=\"12.0\" height=\"10.02\"></span>. So, now we have rewritten original the function we need to approximate as, <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/960fb55b631b92363e66dab5435c9515eeeb3afa0bd5427c5d3391376536ce6c.svg\" alt=\"Typeset mathematical expression\" width=\"108.77\" height=\"10.02\"></span>.</p><p><strong>Ok, so what benefit does this give?  🤷🏻</strong></p><p>By reformulating, we saw was that, <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/83b116716e47411f1f1fbd5f322c8c4ffa04964c4591199d09ce96a740566535.svg\" alt=\"Typeset mathematical expression\" width=\"12.53\" height=\"10.02\"></span> was split into an addition of a function <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/66e75a1dedfe2609c4ae1067eb296cb00e8e9515b2533a5181a3bba32d4b22a5.svg\" alt=\"Typeset mathematical expression\" width=\"12.0\" height=\"10.02\"></span> with the input. In the degradation problem that we saw earlier, the issue was that learning the identity function was hard. However with this residual learning reformulation, it should be easy for the optimizer to drive the weights of the layers such that <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/66e75a1dedfe2609c4ae1067eb296cb00e8e9515b2533a5181a3bba32d4b22a5.svg\" alt=\"Typeset mathematical expression\" width=\"12.0\" height=\"10.02\"></span> becomes a zero mapping. Otherwise, deeper networks should have learnt the identity function in their added layers, giving similar accuracies. In this way, we are left with <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/aab2230694028ee377ca0b0846b8c0faff9178defb2d6557a2c5c20b36271385.svg\" alt=\"Typeset mathematical expression\" width=\"57.29\" height=\"10.02\"></span> which is the identity mapping.</p><p>—</p><figure><p><sup><a href=\"#feed-note-3\">[Note 3]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/eda2a06358c968f7496fecdb126c2b98ceef3ef5020d1ded48523066d8b25b10.png\"></figure><section role=\"doc-bibliography\"><h2>References</h2><ol><li id=\"ref-resnet\" data-paper-title=\"Deep Residual Learning for Image Recognition\"><a href=\"https://arxiv.org/abs/1512.03385\">“Deep Residual Learning for Image Recognition.”</a></li><li id=\"ref-gradient\" data-paper-title=\"The fundamental deep learning problem\"><a href=\"https://people.idsia.ch/~juergen/fundamentaldeeplearningproblem.html\">“The fundamental deep learning problem.”</a></li><li id=\"ref-batchnorm\" data-paper-title=\"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift\"><a href=\"https://arxiv.org/abs/1502.03167\">“Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.”</a></li></ol></section><footer>Unless otherwise noted, the original text of this post is licensed under the <a href=\"https://creativecommons.org/licenses/by/4.0/\">Creative Commons Attribution 4.0 International License</a>. Separately credited material may have different rights. <a href=\"https://new.ramith.fyi/license/\">Licensing details</a>.</footer><h2>Notes</h2><ol><li id=\"feed-note-1\">Figure 1: Credit : Kaiming He, Xiangyu Zhang, Shaoqing Ren, &amp; Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016 [Slides, Video]</li><li id=\"feed-note-2\">Figure 2: Credit : Kaiming He, Xiangyu Zhang, Shaoqing Ren, &amp; Jian Sun. “Deep Residual Learning for Image Recognition”. CVPR 2016 [Slides, Video]</li><li id=\"feed-note-3\">Figure 3: Residual learning block. Source: He et al., “Deep Residual Learning for Image Recognition” <a href=\"#ref-resnet\" role=\"doc-biblioref\">[1]</a>.</li></ol>",
      "date_published": "2022-01-26T00:00:00Z",
      "authors": [
        {
          "name": "Ramith Hettiarachchi",
          "url": "https://new.ramith.fyi"
        }
      ],
      "language": "en",
      "tags": [
        "scholarly"
      ]
    },
    {
      "id": "https://new.ramith.fyi/posts/2022-01-19-imagenet-classification-with-deep-convolutional-neural-networks/",
      "url": "https://new.ramith.fyi/posts/2022-01-19-imagenet-classification-with-deep-convolutional-neural-networks/",
      "title": "ImageNet classification with deep convolutional neural networks (AlexNet)",
      "content_html": "<h2>ImageNet classification with deep convolutional neural networks (AlexNet)</h2><div data-title=\"ImageNet classification with deep convolutional neural networks (AlexNet)\" data-slug=\"2022-01-19-imagenet-classification-with-deep-convolutional-neural-networks\" data-authors=\"Ramith Hettiarachchi\" data-date-published=\"2022-01-19\" data-language=\"en\" data-status=\"published\" data-license=\"CC-BY-4.0\" data-collections=\"scholarly\" data-doi=\"\"><p>By Ramith Hettiarachchi ⋅ Published <time datetime=\"2022-01-19\">2022-01-19</time></p></div><p>The AlexNet paper by Krizhevsky et al. <a href=\"#ref-alexnet\" role=\"doc-biblioref\">[1]</a> was published in 2012. It is a highly influential paper in computer vision which showed that deep networks, along with efficient utilization of GPUs for training, can help build better models. They achieved top-1 and top-5 test set error rates of 37.5% and 17.0%, compared with 47.1% and 28.2% for the previous best results in ILSVRC 2010. These are reductions of 9.6 and 11.2 percentage points, respectively.</p><h3>Introduction</h3><p>Before 2008, the computer vision researchers mostly evaluated their methods on datasets with <strong>tens of thousands of images.</strong> Thanks to the ImageNet dataset <a href=\"#ref-imagenet\" role=\"doc-biblioref\">[2]</a>, which was published in 2009, researchers got the opportunity to move from small scale datasets such as <em>MNIST, CIFAR-10/100</em> to a much larger dataset which has over <strong>15 million labeled images</strong> from more than 22,000 categories. The AlexNet paper evaluates their method for ILSVRC-2010 &amp; 2012 datasets.</p><p>The ImageNet Large-Scale Visual Recognition Challenge (ILSVRC) <a href=\"#ref-ilsvrc\" role=\"doc-biblioref\">[3]</a>, which started in 2010 <strong>focuses on a subset of ImageNet</strong> which has 1.2 million training images, 50,000 validation images, and 150,000 testing images.</p><h3>Architecture</h3><p>Key characteristics of the AlexNet architecture : 8 learned layers (5 Conv, 3 Fully-connected)<br>The architecture has 60 million parameters.</p><figure><p><sup><a href=\"#feed-note-1\">[Note 1]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/447759c0d480c13b6b0fa3f782ef00cc5c65dc83c00def02e489026492732b35.png\"></figure><h4>1. ReLU nonlinearity</h4><p>Inspired by Nair and Hinton’s work <a href=\"#ref-relu\" role=\"doc-biblioref\">[4]</a>, authors have utilized the ReLU non-linearity and have found that it makes the training process several times faster.</p><figure><p><sup><a href=\"#feed-note-2\">[Note 2]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/d8162e7a4b403acca34589ba9d5f9719e20f692dae97ca890690532960e06659.png\"></figure><h4>2. Multiple GPU Training</h4><p>Authors employ a cross-GPU parallelisation approach to train the AlexNet and the communication between GPUs happen only when it is required by the certain layers. They utilize 2x GTX 580 GPUs. Authors mention that this parallelism architecture is similar to “Columnar” CNN by Ciresan et al. <a href=\"#ref-multicolumn\" role=\"doc-biblioref\">[5]</a>.</p><figure><p><sup><a href=\"#feed-note-3\">[Note 3]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/bd3bbdfac034c4e11175dd57b96c519a8a0b5419cf369625e7e7d9fd700c321e.png\"></figure><h4>3. Local Response Normalization</h4><p>Krizhevsky et al. has found that their local normalization methods helps achieve better generalization. For this procedure, they sum over <span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/67bea37ea4704b501be7fe53d57c02ede9d6d9a19b6c327330e5ee8c5d0c3e3c.svg\" alt=\"Typeset mathematical expression\" width=\"8.8\" height=\"10.02\"></span> adjacent kernel maps at the same spatial location as shown in the equation below. This normalization has been applied in certain layers of AlexNet. <strong><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/7d9273b2d6fe06e5e66aa8d754f5e4e7ced443ddc47f1426127bceae8dc439d7.svg\" alt=\"Typeset mathematical expression\" width=\"13.33\" height=\"10.02\"></span></strong> = total number of kernels<br><strong><span role=\"math\"><img src=\"https://new.ramith.fyi/feed-assets/56cf8a01c1d9ccf6f87a1bacbd85cb02bc130a0e83b1cc4f21735d155d3f082c.svg\" alt=\"Typeset mathematical expression\" width=\"33.55\" height=\"10.13\"></span></strong> = response-normalized neuron activity</p><figure><p><sup><a href=\"#feed-note-4\">[Note 4]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/d470eceeefb022d86e6f8d3ebf8b56c20c56c9ace04354d82c945658504b78de.png\"></figure><h4>4. Overlapping Pooling</h4><p>By overlapping pooling scheme has reduced top-1 and top-5 error rates by 0.4% and 0.3%. Furthermore, the authors have observed that models with overlapping pooling are slightly more difficult to overfit.</p><h3>Reducing Overfitting</h3><p>Authors have utilized <strong>1) data augmentations</strong> , <strong>2) dropout</strong> to reduce overfitting. Data augmentation has been done in two ways.</p><p><strong>1)</strong> Generating image translations and horizontal reflections</p><p><strong>2)</strong> Altering RGB channel intensities in training images</p><p>Authors cite their early work on dropout <a href=\"#ref-dropout\" role=\"doc-biblioref\">[6]</a> and mention that it has reduced overfitting substantially while roughly doubling the number of training iterations needed to converge.</p><section role=\"doc-bibliography\"><h2>References</h2><ol><li id=\"ref-alexnet\" data-paper-title=\"ImageNet Classification with Deep Convolutional Neural Networks\">Krizhevsky et al. (2012). <a href=\"https://proceedings.neurips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html\">“ImageNet Classification with Deep Convolutional Neural Networks.”</a></li><li id=\"ref-imagenet\" data-paper-title=\"ImageNet: A Large-Scale Hierarchical Image Database\">Deng et al. (2009). <a href=\"https://image-net.org/static_files/papers/imagenet_cvpr09.pdf\">“ImageNet: A Large-Scale Hierarchical Image Database.”</a></li><li id=\"ref-ilsvrc\" data-paper-title=\"ImageNet Large Scale Visual Recognition Challenge\">Russakovsky et al. (2014). <a href=\"https://arxiv.org/abs/1409.0575\">“ImageNet Large Scale Visual Recognition Challenge.”</a></li><li id=\"ref-relu\" data-paper-title=\"Rectified Linear Units Improve Restricted Boltzmann Machines\">Nair and Hinton (2010). <a href=\"https://www.cs.toronto.edu/~hinton/absps/reluICML.pdf\">“Rectified Linear Units Improve Restricted Boltzmann Machines.”</a></li><li id=\"ref-multicolumn\" data-paper-title=\"Multi-column Deep Neural Networks for Image Classification\">Ciresan et al. (2012). <a href=\"https://arxiv.org/abs/1202.2745\">“Multi-column Deep Neural Networks for Image Classification.”</a></li><li id=\"ref-dropout\" data-paper-title=\"Improving neural networks by preventing co-adaptation of feature detectors\">Hinton et al. (2012). <a href=\"https://arxiv.org/abs/1207.0580\">“Improving neural networks by preventing co-adaptation of feature detectors.”</a></li></ol></section><footer>Unless otherwise noted, the original text of this post is licensed under the <a href=\"https://creativecommons.org/licenses/by/4.0/\">Creative Commons Attribution 4.0 International License</a>. Separately credited material may have different rights. <a href=\"https://new.ramith.fyi/license/\">Licensing details</a>.</footer><h2>Notes</h2><ol><li id=\"feed-note-1\">Figure 1:<span>&#x20;</span></li><li id=\"feed-note-2\">Figure 2:<span>&#x20;</span></li><li id=\"feed-note-3\">Figure 3: Just to get an idea on how GPUs now and then (2012) compare with each other. Source : GadgetVersus</li><li id=\"feed-note-4\">Figure 4: Local reponse normalization has reduced top-1 and top-5 error rates by 1.4% and 1.2%.</li></ol>",
      "date_published": "2022-01-19T00:00:00Z",
      "authors": [
        {
          "name": "Ramith Hettiarachchi",
          "url": "https://new.ramith.fyi"
        }
      ],
      "language": "en",
      "tags": [
        "scholarly"
      ]
    },
    {
      "id": "https://new.ramith.fyi/posts/2022-01-05-jax-001-a-closer-look-at-its-background/",
      "url": "https://new.ramith.fyi/posts/2022-01-05-jax-001-a-closer-look-at-its-background/",
      "title": "JAX 001 - A closer look at its background",
      "content_html": "<h2>JAX 001 - A closer look at its background</h2><div data-title=\"JAX 001 - A closer look at its background\" data-slug=\"2022-01-05-jax-001-a-closer-look-at-its-background\" data-authors=\"Ramith Hettiarachchi\" data-date-published=\"2022-01-05\" data-language=\"en\" data-status=\"published\" data-license=\"CC-BY-4.0\" data-collections=\"scholarly\" data-doi=\"\"><p>By Ramith Hettiarachchi ⋅ Published <time datetime=\"2022-01-05\">2022-01-05</time></p></div><p>I recently wanted to get into parallel programming so that I can optimize one of the optics simulation projects I was working on (it was built without much focus on distributed training). So with all the buzz going around with JAX and some cool applications I found (listed below), I thought to give it a try.</p><ul><li>Protein Structure Prediction -<a href=\"https://github.com/deepmind/alphafold\">AlphaFold 2</a> <a href=\"#ref-alphafold\" role=\"doc-biblioref\">[1]</a></li><li>Differentiable, Hardware Accelerated, Molecular Dynamics -<a href=\"https://github.com/google/jax-md\">JAX-MD</a> <a href=\"#ref-jaxmd\" role=\"doc-biblioref\">[2]</a></li><li>Massively parallel rigid-body physics simulation -<a href=\"https://github.com/google/brax\">Brax</a> <a href=\"#ref-brax\" role=\"doc-biblioref\">[3]</a></li><li>Chemical Modelling -<a href=\"https://github.com/deepchem/jaxchem\">JAXChem</a> <a href=\"#ref-jaxchem\" role=\"doc-biblioref\">[4]</a></li><li>Computational Fluid Dynamics -<a href=\"https://github.com/google/jax-cfd\">JAX-CFD</a> <a href=\"#ref-jaxcfd\" role=\"doc-biblioref\">[5]</a></li><li>Differentiable Cosmology -<a href=\"https://github.com/DifferentiableUniverseInitiative/jax_cosmo\">JAX-Cosmo</a> <a href=\"#ref-jaxcosmo\" role=\"doc-biblioref\">[6]</a></li></ul><h3>Why JAX?</h3><p>With all the established libraries such as PyTorch and Tensorflow, why is there a requirement for this new library in the first place? Let’s find out.</p><p>We all know that increasing FLOPS (floating point operations per second) is a huge deal in machine learning to train models efficiently. JAX aims to help this goal by enabling researchers to write python programs which are automatically <strong>compiled</strong> and <strong>scaled</strong> to utilize accelerators (GPUs/TPUs). Often it is<a href=\"https://twitter.com/ChrSzegedy/status/1400704509240692743\">hard</a> to write optimized code in python to leverage the potential of hardware accelerators. JAX aims to keep a balance between research-friendly programming experience vs hardware acceleration.</p><p>To do so, JAX aims to accelerate <em>pure-and-statically-composed (PSC) subroutines</em>. It uses a <strong>just-in-time (JIT)</strong> compiler that traces these routines. To learn more about PSC subroutines, see the <a href=\"https://www.youtube.com/watch?v=iDxJxIyzSiM\"><em>NeurIPS 2020: JAX Ecosystem Meetup</em> video</a>. The compilation process monitors the program once; the name JAX was described as <em>“Just After eXecution”</em> in the <a href=\"https://cs.stanford.edu/~rfrostig/pubs/jax-mlsys2018.pdf\">JAX paper</a> <a href=\"#ref-resource1\" role=\"doc-biblioref\">[8]</a>.</p><h3>Key features of JAX</h3><p>JAX addresses several limitations in numpy. Therefore, it presents,</p><ul><li>Lightweight NumPy-like API for array-based computing</li><li>Composable function transformations<br></li></ul><p>(autodiff, JIT compilation, vectorization, parallelization)</p><ul><li>Execute on CPU, GPU, or TPU without changing your code.</li></ul><p>&gt; JAX is Autograd and XLA, brought together for high-performance numerical computing and machine learning research. It provides composable transformations of Python+NumPy programs: differentiate, vectorize, parallelize, Just-In-Time compile to GPU/TPU, and more.</p><figure><p><sup><a href=\"#feed-note-1\">[Note 1]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/5bd72b3a3f200197caf0f27c9a0ae624717ecb38339ce9121694bf4ef058ef19.png\"></figure><h3>Getting familiar with related concepts and history</h3><p>There were a few loosly defined terms in my mind when I first read through the documentation. So I thought to look into those and have an idea about them.</p><p><em><strong>Autograd -</strong></em> Autograd is an example of an <em>automatic differentiation</em> library<a href=\"https://github.com/HIPS/autograd\">released in 2014</a> . It is a lightweight tool to automatically differentiate native python and numpy code. But it doesn’t focus on hardware accelerators such GPU/TPUs. Therefore, since 2018, the <strong>main developers of Autograd are focusing on JAX</strong> which has more functionalities. Dougal, one of the main authors of autograd,<a href=\"https://youtu.be/5XRphWnFkNM?t=376\">calls</a> JAX as the second generation of the project which started with Autograd.</p><p><strong><em>“XLA</em></strong> <strong>-</strong> <strong>(Accelerated Linear Algebra)</strong> is a compiler-based linear algebra execution engine. It is the backend that powers machine learning frameworks such as TensorFlow and JAX at Google, on a variety of devices including CPUs, GPUs, and TPUs.” - (to<a href=\"https://developers.googleblog.com/2017/03/xla-tensorflow-compiled.html\">read more</a> )</p><p><em><strong>JIT</strong> <strong>- (Just in time)</strong> compilation: This enables</em> compilation of a JAX Python function so it can be executed efficiently in XLA.</p><p>Digging into these concepts made me realize that while the libraries we use makes it really easy to build end-to-end differentiable models, under the hood, these tools have been years-long effort by amazing teams with cool design choices.</p><p>In the next writeup I’m going to share my experience with device parallelization (<code>pmap</code>) and automatic vectorization (<code>vmap</code>) aspects.</p><p>—</p><section role=\"doc-bibliography\"><h2>References</h2><ol><li id=\"ref-alphafold\" data-paper-title=\"Highly accurate protein structure prediction with AlphaFold\">Jumper et al. (2021). <a href=\"https://doi.org/10.1038/s41586-021-03819-2\">“Highly accurate protein structure prediction with AlphaFold.”</a></li><li id=\"ref-jaxmd\" data-paper-title=\"JAX MD: A Framework for Differentiable Physics\">Schoenholz and Cubuk (2020). <a href=\"https://papers.neurips.cc/paper_files/paper/2020/hash/83d3d4b6c9579515e1679aca8cbc8033-Abstract.html\">“JAX MD: A Framework for Differentiable Physics.”</a></li><li id=\"ref-brax\" data-paper-title=\"Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation\">Freeman et al. (2021). <a href=\"https://arxiv.org/abs/2106.13281\">“Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation.”</a></li><li id=\"ref-jaxchem\" data-paper-title=\"JAXChem: JAX-based chemical modeling software\">JAXChem contributors <a href=\"https://github.com/deepchem/jaxchem\">“JAXChem: JAX-based chemical modeling software.”</a></li><li id=\"ref-jaxcfd\" data-paper-title=\"Machine learning–accelerated computational fluid dynamics\">Kochkov et al. (2021). <a href=\"https://doi.org/10.1073/pnas.2101784118\">“Machine learning–accelerated computational fluid dynamics.”</a></li><li id=\"ref-jaxcosmo\" data-paper-title=\"JAX-Cosmo: A differentiable cosmology library in JAX\">JAX-Cosmo contributors <a href=\"https://github.com/DifferentiableUniverseInitiative/jax_cosmo\">“JAX-Cosmo: A differentiable cosmology library in JAX.”</a></li><li id=\"ref-resource0\" data-paper-title=\"JAX documentation\"><a href=\"https://jax.readthedocs.io/en/latest/index.html\">“JAX documentation.”</a></li><li id=\"ref-resource1\" data-paper-title=\"Compiling machine learning programs via high-level tracing\">Frostig, Johnson, and Leary (2018). <a href=\"https://cs.stanford.edu/~rfrostig/pubs/jax-mlsys2018.pdf\">“Compiling machine learning programs via high-level tracing.”</a></li><li id=\"ref-resource2\" data-paper-title=\"JAX: accelerated ML research — video\"><a href=\"https://youtu.be/mVf3HJ6gNDc\">“JAX: accelerated ML research — video.”</a></li><li id=\"ref-resource3\" data-paper-title=\"JAX: accelerated ML research — slides\"><a href=\"https://program-transformations.github.io/slides/NeurIPS_workshop_JAX_talk.pdf\">“JAX: accelerated ML research — slides.”</a></li><li id=\"ref-resource4\" data-paper-title=\"JAX at DeepMind — video\"><a href=\"https://www.youtube.com/watch?v=iDxJxIyzSiM\">“JAX at DeepMind — video.”</a></li><li id=\"ref-resource5\" data-paper-title=\"JAX at DeepMind — slides\"><a href=\"https://storage.googleapis.com/deepmind-media/Jax/NeurIPS%20outreach%20session.pdf\">“JAX at DeepMind — slides.”</a></li><li id=\"ref-resource6\" data-paper-title=\"Lecture 6: Automatic Differentiation\"><a href=\"https://www.cs.toronto.edu/~rgrosse/courses/csc421_2019/readings/L06%20Automatic%20Differentiation.pdf\">“Lecture 6: Automatic Differentiation.”</a></li><li id=\"ref-resource7\" data-paper-title=\"Cloud TPU Research Program\"><a href=\"https://sites.research.google/trc/about/\">“Cloud TPU Research Program.”</a></li></ol></section><footer>Unless otherwise noted, the original text of this post is licensed under the <a href=\"https://creativecommons.org/licenses/by/4.0/\">Creative Commons Attribution 4.0 International License</a>. Separately credited material may have different rights. <a href=\"https://new.ramith.fyi/license/\">Licensing details</a>.</footer><h2>Notes</h2><ol><li id=\"feed-note-1\">Figure 1: Credit: {mattjj, frostig, leary, dougalm, phawkins, skyewm, jekbradbury, necula} @google.com. Full set of slides: <a href=\"#ref-resource3\" role=\"doc-biblioref\">[10]</a>.</li></ol>",
      "date_published": "2022-01-05T00:00:00Z",
      "authors": [
        {
          "name": "Ramith Hettiarachchi",
          "url": "https://new.ramith.fyi"
        }
      ],
      "language": "en",
      "tags": [
        "scholarly"
      ]
    },
    {
      "id": "https://new.ramith.fyi/posts/2021-01-09-setting-up-raspberry-pi-4-with-ubuntu-20-04-ros-intel-realsense/",
      "url": "https://new.ramith.fyi/posts/2021-01-09-setting-up-raspberry-pi-4-with-ubuntu-20-04-ros-intel-realsense/",
      "title": "Getting the installation Right! Raspberry Pi 4 with Ubuntu 20.04 + ROS Noetic + Intel RealSense",
      "content_html": "<h2>Getting the installation Right! Raspberry Pi 4 with Ubuntu 20.04 + ROS Noetic + Intel RealSense</h2><div data-title=\"Getting the installation Right! Raspberry Pi 4 with Ubuntu 20.04 + ROS Noetic + Intel RealSense\" data-slug=\"2021-01-09-setting-up-raspberry-pi-4-with-ubuntu-20-04-ros-intel-realsense\" data-authors=\"Ramith Hettiarachchi\" data-date-published=\"2021-01-09\" data-language=\"en\" data-status=\"published\" data-license=\"CC-BY-4.0\" data-collections=\"scholarly\" data-doi=\"\"><p>By Ramith Hettiarachchi ⋅ Published <time datetime=\"2021-01-09\">2021-01-09</time></p></div><p><em>Historical guide: these steps describe my January 2021 setup with Ubuntu 20.04 and ROS Noetic.</em></p><p>Here I am after lot of iterations trying to install Ubuntu on Raspberry Pi 4(RPi4). Initially I made certain mistakes when selecting the correct ubuntu version &amp; Desktop.</p><p>The steps below are which I followed to finalize my Rpi Installation with Ubuntu.</p><h3>Ubuntu on RPi - Confusions?</h3><p>My goal was to have a <strong>stable ROS installation</strong> on <strong>ubuntu on RPI4</strong> . When considering the dependencies, I realized that RPI officially supports Ubuntu 20.04 &amp; 20.10. And as of now, the ROS version that matches with one of those ubuntu versions (Ubuntu 20.04) is ROS Noetic.</p><p>Now in ubuntu’s official download page (For RPi), they have listed the following versions. However, Ubuntu 20.04 Desktop version is not listed here. Therefore, the goal should be to install Ubuntu Server 20.04 and then install a desktop of our choice.</p><figure><p><sup><a href=\"#feed-note-1\">[Note 1]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/80a7eae8421c19b0aaf04a94c9190eb785292206c7d5ade2ed851532e3987928.png\"></figure><h4>How this guide is organized</h4><ol><li>Installing Ubuntu Server 20.04.1<br></li></ol><ul><li>Setting up SD card (through RPi Imager)<br></li><li>Editing <strong>network-config</strong> file =&gt; connect to network</li></ul><ol><li>Installing the Desktop for Ubuntu Server</li><li>Trying out screen sharing<br></li></ol><ul><li>Connect remotely to view desktop</li></ul><ol><li>Installing ROS Noetic</li><li>Installing Realsense libraries for Ubuntu 20.04</li></ol><h4>1. Installing Ubuntu Server 20.04.1</h4><p>To install ubuntu server follow the official guide. When editing the file in microSD, double check whether you have edited it properly with correct wifi credentials</p><p><a href=\"https://ubuntu.com/tutorials/how-to-install-ubuntu-on-your-raspberry-pi#1-overview\"><a href=\"https://ubuntu.com/tutorials/how-to-install-ubuntu-on-your-raspberry-pi#1-overview\">https://ubuntu.com/tutorials/how-to-install-ubuntu-on-your-raspberry-pi#1-overview</a></a></p><h4>2. Installing the Desktop for Ubuntu Server</h4><p>This is a tricky part. After searching for “how to install desktop on Ubuntu Server”, I tried the command ❌ <code>sudo apt-get install ubuntu-desktop</code>, which led to many errors in the desktop environment, such as Wi-Fi connections not displaying.</p><p>Therefore, we will be using a tool called<a href=\"https://github.com/wimpysworld/desktopify\"><em><strong>Desktopify</strong></em></a> <em><strong>.</strong></em> After you have installed ubuntu server, follow these steps to install the Desktop</p><pre><code>git clone https://github.com/wimpysworld/desktopify.git\ncd desktopify\nsudo ./desktopify --de ubuntu</code></pre><p>This will take some time to install. This script will take care of all the dependencies you need, and will install a clean desktop environment</p><figure><p><sup><a href=\"#feed-note-2\">[Note 2]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/a8b8d10db2a0c498a0bb11e7f7e3e1911773a1b85aedbbda6a544b8c977f0210.png\"></figure><h4>3. Trying out screensharing</h4><p>Please follow the guide<a href=\"https://ubuntuhandbook.org/index.php/2020/07/remote-desktop-sharing-ubuntu-20-04/\">here</a> to enable screensharing so that you can view the RPi’s desktop remotely from your Laptop.</p><p><strong>NOTE: To connect from your computer you will first need a monitor which is plugged to the RPI4.</strong> <em>If you want to <strong>access the screen remotely even when a monitor is not attached</strong> , Check the <strong>very end</strong> of this post.</em></p><figure><p><sup><a href=\"#feed-note-3\">[Note 3]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/f35e7d51b7182acb701c145280f4485022183b48b5835c82655f4f21fd2b62b4.png\"></figure><h4>4. Installing ROS Noetic</h4><p>You can follow these steps, to install ROS Noetic</p><p><a href=\"http://wiki.ros.org/noetic/Installation/Ubuntu\"><a href=\"http://wiki.ros.org/noetic/Installation/Ubuntu\">http://wiki.ros.org/noetic/Installation/Ubuntu</a></a></p><h4>5. Installing Realsense libraries</h4><p>My friend Shalutha suggested these steps to install the required dependencies of Realsense.</p><p><a href=\"https://answers.ros.org/question/363889/intel-realsens-on-ubuntu-2004-ros-noetic-installation-desription/\">RealSense installation instructions on ROS Answers</a>.</p><p>You can copy paste these steps to an .sh file and execute it</p><pre><code>nano realsense.sh\n#copy paste the steps listed in the link to this file\nchmod +x realsense.sh\n./realsense.sh</code></pre><figure><p><sup><a href=\"#feed-note-4\">[Note 4]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/15df4c5f540744805124c8e56f24072eff45b142df49943d61c84b0e277c4704.png\"></figure><h4>Congratulations 🥳</h4><p>Now you have a stable ubuntu desktop 20.04 installation!</p><pre><code>/opt/realsense/bin/realsense-viewer</code></pre><figure><p><sup><a href=\"#feed-note-5\">[Note 5]</a></sup></p><img src=\"https://new.ramith.fyi/feed-assets/5766b0ff2050cfcc8e538a15b298498fef145f62a43253b625a1e6faeea4e56e.png\"></figure><h3>Extras</h3><h4>1. Logging in with an SSH key</h4><p>Follow the <a href=\"https://www.atlantic.net/vps-hosting/how-to-set-up-ssh-keys-on-ubuntu-18-04/\">SSH key setup tutorial</a> linked in the original guide.</p><h4>2. Accessing remote screen when a monitor is not attached.</h4><pre><code>sudo apt-get install xserver-xorg-video-dummy</code></pre><pre><code>sudo nano /etc/X11/xorg.conf\n</code></pre><p>Copy the following configuration into that file:</p><pre><code>Section &quot;Device&quot;\n    Identifier &quot;Configured Video Device&quot;\n    Driver &quot;dummy&quot;\nEndSection\n\nSection &quot;Monitor&quot;\n    Identifier &quot;Configured Monitor&quot;\n    HorizSync 31.5-48.5\n    VertRefresh 50-70\nEndSection\n\nSection &quot;Screen&quot;\n    Identifier &quot;Default Screen&quot;\n    Monitor &quot;Configured Monitor&quot;\n    Device &quot;Configured Video Device&quot;\n    DefaultDepth 24\n    SubSection &quot;Display&quot;\n        Depth 24\n        Modes &quot;1024x800&quot;\n    EndSubSection\nEndSection</code></pre><p>Now reboot the RPi<code>sudo reboot</code>. Then you can connect from your laptop even if there is no monitor attached to the RPi.</p><p>(source:<a href=\"https://askubuntu.com/questions/453109/add-fake-display-when-no-monitor-is-plugged-in\"><a href=\"https://askubuntu.com/questions/453109/add-fake-display-when-no-monitor-is-plugged-in\">https://askubuntu.com/questions/453109/add-fake-display-when-no-monitor-is-plugged-in</a></a> )</p><h3>Final outcome of setting up 🎉</h3><div><p><a href=\"https://www.youtube.com/watch?v=CgjEviTkvpA\">Watch the original setup demo on YouTube.</a></p></div><footer>Unless otherwise noted, the original text of this post is licensed under the <a href=\"https://creativecommons.org/licenses/by/4.0/\">Creative Commons Attribution 4.0 International License</a>. Separately credited material may have different rights. <a href=\"https://new.ramith.fyi/license/\">Licensing details</a>.</footer><h2>Notes</h2><ol><li id=\"feed-note-1\">Figure 1: source: <a href=\"https://ubuntu.com/download/raspberry-pi\">https://ubuntu.com/download/raspberry-pi</a></li><li id=\"feed-note-2\">Figure 2: After Desktopifying Ubuntu Server 20.04.1</li><li id=\"feed-note-3\">Figure 3:<span>&#x20;</span></li><li id=\"feed-note-4\">Figure 4: (It will take some time to install)</li><li id=\"feed-note-5\">Figure 5: (Screenshot taken from the plugged monitor)</li></ol>",
      "date_published": "2021-01-09T00:00:00Z",
      "authors": [
        {
          "name": "Ramith Hettiarachchi",
          "url": "https://new.ramith.fyi"
        }
      ],
      "language": "en",
      "tags": [
        "scholarly"
      ]
    }
  ]
}
