# ReLU NTNU - Full content > ReLU NTNU is a student-run applied machine learning organization at the Norwegian University of Science and Technology in Trondheim. Total published posts: 5 Total listed projects: 27 --- ## Published Writing ### The case for AI interpretability in the agentic era URL: /blog/interpretability-agentic-era/ Date: 2026-03-13 | Author: Knut Opland Moen Øyvind Tafjord, Research Scientist in Interpretability at Google DeepMind, on emergence, the agentic shift, the 'golden window' of Chain-of-Thought, and why we need tools to dig as deep as we want into the systems we deploy. As artificial intelligence transitions from passive chatbots to autonomous agents capable of executing complex tasks, the "black box" problem is no longer just an academic curiosity — it is a critical safety challenge. After his guest lecture at ReLU, we asked Øyvind Tafjord, Research Scientist in Interpretability at Google DeepMind, a few questions regarding how we can maintain trust in systems that are becoming increasingly complex. The field of Artificial Intelligence is currently undergoing a structural change. We are moving from the era of Generative AI (systems that predict text or pixels) to the era of Agentic AI: systems that plan, reason, use tools, and execute decisions in the real world. While this promises unprecedented utility, it brings a distinct crisis of transparency. When an AI is managing financial audits or medical diagnostics, being "right" is no longer enough. We need to know why it is right. ## 1. Scale surprises To understand where we are going, we must appreciate the surprise of how we got here. For decades, the assumption was that reasoning required complex, symbolic rules programmed by humans. The reality of Large Language Models (LLMs) proved otherwise. As Tafjord notes, the field was caught off guard by what simple scaling could achieve: > "Language models have existed for a long time, but it was surprising that with a sufficiently rich model (e.g., transformer) and enough data, the models can indirectly learn everything possible about the world, from language structure to facts and reasoning patterns. BERT and GPT-2 were already shocking in what they could perform, but are still very primitive compared to the newest models." This phenomenon is known in the research literature as *emergence*. We now know that with sufficient scale, models don't just memorize; they develop internal representations that support logic, arithmetic, and even theory-of-mind‑like behavior. However, these capabilities often emerge within a "black box," leaving researchers to reverse-engineer how the model actually learned them. As these emergent capabilities grow, the focus shifts from *what* the model can do to *how* it actually executes those tasks. ## 2. The agentic shift The stakes of the black box rise with autonomous agents. Unlike a chatbot that offers a suggestion, an agent acts. It might access a database, execute code, or control an API. In this context, the final output is insufficient for trust. If an agent denies a loan application or recommends a specific medical treatment, we require an audit trail of its logic. Tafjord emphasizes that future systems must be built for interrogation: > "Generally, I feel it is important that we set up our AI systems so that we can have an understanding of why important decisions are taken. Not necessarily that we understand the 'innards' of the models… but that we can follow the reasoning that leads to the answer. Preferably in an interactive way, so that we can dig as deep as we want, where needed." Kim et al. (2025) suggest that this interactive "agentic interpretability" serves as a necessary stress test for autonomous systems. They compare this dynamic to a cross-examination: while a model might easily fabricate a single justification for a loan denial, maintaining a consistent facade of logic across a multi-turn interrogation requires significant "cognitive load." By forcing the agent to explain its reasoning over a long dialogue, and potentially performing "open-model surgery" to verify those claims against internal states, we make it much harder for an agent to hide deceptive logic or incompetence. To satisfy the regulatory need for documentation, this conversation concludes not with silence, but with a generated "meeting note" — a consensus report that solidifies the dialogue into a verifiable audit trail. Moving on, if we are to interrogate these agents effectively, we must look at the specific methods currently available to look into their "thoughts." ## 3. The "golden window" of Chain-of-Thought So, how do we peek inside? Currently, we are in a fortunate, but perhaps temporary, "golden window" of transparency. State-of-the-art models utilize Chain-of-Thought (CoT) reasoning. They "think" by generating tokens in human language before producing a final answer. This allows us to monitor their logic, catch hallucinations, and understand their intent. However, Tafjord warns that this reliance on human language for machine thinking might not last: > "It is quite practical that many of the current models operate with ordinary language as a form of communication… This makes it possible to get a sort of superficial understanding of how they reason. This may disappear eventually, if models use internal, more effective representations to 'think' or communicate with each other — that can become a problem." Recent research supports this concern. Concepts like reward hacking and steganography suggest that as models are optimized for pure efficiency, they may learn to "hide" information in high-dimensional vectors that humans cannot read. If an AI creates its own shorthand for logic that is 100x more efficient than English, we lose our ability to audit its thoughts. ## 4. Understanding the AI brain If the models stop speaking our language, how do we understand them? This is where the field shifts from "behavioral psychology" (watching what the model says) to "digital neuroscience" (mapping the neurons). Tafjord highlights this duality: > "When it comes to 'understanding' such systems, there are many approaches (analogous to how one can study animals/humans through how they behave, or try to model from more fundamental mechanisms for what happens inside them)." The "fundamental mechanisms" approach is known as **Mechanistic Interpretability**. Researchers can in a way treat neural networks like biological organisms. Using tools like Sparse Autoencoders, they are decomposing the "neural soup" of a model into distinct features, identifying specific circuits responsible for deception, poetic structure, or coding ability. Recent breakthroughs in monosemanticity allow us to find the specific "direction" in a model's "brain" that activates when it is thinking about, for example, the Golden Gate Bridge or a specific Python function. This is the future Tafjord alludes to: moving beyond "superficial understanding" to a precise wiring diagram of machine intelligence. ## 5. The optimism of exploration Despite the daunting complexity, there is reason for optimism. The field is not stagnant; it is exploding with creative solutions. We are seeing the rise of **Mixture of Experts (MoE)** architectures, which compartmentalize knowledge, and **Retrieval-Augmented Generation (RAG)**, which grounds models in external facts. Tafjord finds this rapid iteration fascinating: > "The development of language models has continued with (accelerating?) progression… Most ideas one sees published don't lead very far, but because so many different things are being attempted, both in academia and industry, we constantly find things that actually make things significantly better. It is quite fascinating from a bird's eye view." This "bird's eye view" reveals a landscape where failures are frequent but necessary, driving the incremental breakthroughs that redefine what is possible. **Mixture of Experts (MoE)** is a model architecture that divides knowledge into specialized sub-networks ("experts"), where only the relevant expert is activated for a given task. This makes large models more efficient and their knowledge more compartmentalized. **Retrieval-Augmented Generation (RAG)** is a technique that supplements a model's built-in knowledge by letting it fetch relevant information from an external database at runtime, grounding its answers in verifiable, up-to-date facts rather than relying solely on what it learned during training. ## Conclusion: trust through transparency As AI systems are deployed across high-stakes domains such as healthcare, law, and financial services, technical performance alone is an insufficient benchmark for trustworthiness. It is no longer enough for a model to be smart; it must be intelligible. The goal is to ensure that as these systems scale, our understanding scales with them. We must build the tools today that allow us to "dig as deep as we want" tomorrow, ensuring that artificial intelligence remains a tool we can not only use, but truly understand. --- ### More than a GPT-wrapper: why top-tier ML engineers master the geometry of math URL: /blog/sondre-stanford/ Date: 2026-03-04 | Author: Knut Opland Moen ReLU alumnus Sondre Rogde, now at Stanford's ICME programme, on why the best Quants still love Linear Regression, the geometry behind L1 and L2 regularization, and the math foundations every ML engineer should revisit. ![Sondre Rogde](/blog/sondre/sondre.webp) What distinguishes a "GPT-wrapper" developer from a top-tier Machine Learning Engineer? According to our recent chat with Sondre, a ReLU alumnus, the answer often lies in Linear Algebra. Sondre is currently pursuing a Master's in Computational and Mathematical Engineering (ICME) at Stanford University. While his time at NTNU provided a strong technical foundation, his current program has pushed him deep into the theoretical underpinnings of the algorithms we often take for granted. We caught up with him to discuss the importance of mathematical intuition, why the best Quants in the world still love Linear Regression, and the geometric concepts every data scientist should master. ## The shift: from implementation to "pure math" While many Computer Science programs focus heavily on implementation and software development, Sondre's track at Stanford is distinctively different. It is less about importing libraries and more about proving why those libraries work. "It's a lot of work, and the focus is heavily theoretical," Sondre explains. "Many students here, especially those from France, come from pure math backgrounds. We spend a lot of time on proofs and Linear Algebra. It's a different perspective than the standard Computer Science approach, but it gives you total control over what you are building." ## The academic take: why you should revisit Linear Regression One of the most surprising insights from Sondre's time in Silicon Valley is the industry's reliance on "basic" models. He recounts speaking with the Head of Risk at a major quantitative firm who relies almost exclusively on Linear Regression. Why? Interpretability and control. "When you use a complex neural network, you often lose the ability to explain why the model made a specific prediction," Sondre notes. "In high-stakes fields like finance, if you can't explain the output, nobody will listen to you." But it goes deeper than just "keeping it simple." Sondre highlights that a deep understanding of regularization allows Linear Regression to handle complex, correlated data just as well as black-box models. ### 1. The geometry of regularization (L1 vs. L2) Sondre is currently taking a course taught by Trevor Hastie, a legend in the field and co-author of *The Elements of Statistical Learning*. A key takeaway from Hastie's course is the geometric understanding of norms. "We often just memorize the formulas for L1 (Lasso) and L2 (Ridge) regularization, but understanding their geometry changes how you view the model," Sondre says. - **L1 norm (Lasso):** geometrically, the constraint region looks like a diamond. Because the solution often hits the "corners" of this diamond, L1 regularization drives coefficients to exactly zero. This creates sparse solutions, effectively performing feature selection for you. - **L2 norm (Ridge):** the constraint is a circle. It shrinks coefficients but rarely zeros them out entirely, which is better for handling multicollinearity without removing variables. "I've become a bit in love with regressions because of this," Sondre admits. "When you understand the geometry, you realize that Linear Regression involves projecting data onto a subspace. It's elegant, and you know exactly what is happening to your variables." ### 2. Orthogonality and PCA Another concept Sondre wishes he had utilized more during his projects at NTNU is Principal Component Analysis (PCA) — not just for dimensionality reduction, but for handling correlation. "In real-world projects, like the ones we do in ReLU, you often have variables that are highly correlated," he explains. "If you feed that straight into a model, your coefficients (betas) stop representing the marginal effect of a single variable, because they are entangled with others." By using techniques like PCA, you can transform your features to be orthogonal (statistically independent). This ensures that your model is stable and your interpretations are valid. "Standardizing variables and understanding their correlation structure is much more important than people think," he adds. Sondre specifically reflected on his previous work experience at Norges Bank Investment Management (NBIM) to illustrate this point. "I think I could have used these techniques a lot more in the project I had at NBIM," he says. Working with financial data often means dealing with a massive number of variables that move together. Looking back, he realizes that simply feeding that data into a black-box model isn't optimal. By using PCA to transform those variables so they are orthogonal before modeling, he could have achieved a much cleaner, more robust interpretation of the market drivers. ![Stanford's Main Quad on a sunny afternoon](/blog/sondre/stanford-quad.webp) Caption: A sunny afternoon at Stanford's historic Main Quad (Photo: Sondre Rogde) ## The "AI bubble" vs. deep engineering Living in the Bay Area, Sondre sees the hype cycle up close. He distinguishes between "GPT-wrappers" — startups that simply call an API — and companies doing "deep engineering." "I don't think the AI industry is a bubble for the big players, but for the shallow applications, it might be," he observes. He points to tools like Cursor as an example of AI done right. "The value isn't just in using a language model. The value is in the engineering required to make math agents understand a codebase and perform complex reasoning. That requires a level of mathematical maturity that goes beyond just prompting a model." ## Final advice for ReLU students Sondre's advice to current members is not to shy away from the advanced models, but to ensure the foundation is solid first. "At NTNU, we have the freedom to build really cool, practical things. But my advice is to take the time to understand the data before you throw a Deep Learning model at it. Look at the correlations. Consider a baseline Linear Regression. Understand the geometry of your loss function." "My vision for ReLU was always to bridge the gap between academia and industry," Sondre concludes. "And in the industry, the best engineers are often the ones who master the basics." --- ### From ReLU to Berkeley: Jørgen's journey into the heart of AI URL: /blog/jorgen-berkeley/ Date: 2026-02-13 | Author: Knut Opland Moen A ReLU member on exchange at UC Berkeley reflects on the workload, the night culture, the integration of research into undergraduate life, and what Norwegian students can learn from the American mindset. ![Jørgen at UC Berkeley](/blog/jorgen/jorgen.webp) What happens when you combine a hunger for machine learning, a strong student community at NTNU, and a semester in Silicon Valley? We recently caught up with Jørgen, a ReLU member who is currently on exchange at UC Berkeley. We wanted to hear about his transition from Trondheim to the competitive academic landscape of the U.S., his experience participating in San Francisco hackathons, and how his time in ReLU prepared him for the global stage. Here is a look into Jørgen's journey, the cultural differences he's encountered, and his advice for aspiring AI engineers. ## The ReLU foundation For Jørgen, the journey didn't start in California; it started in Trondheim with a simple desire to understand the buzz around Artificial Intelligence. "I initially wanted to join ReLU because I understood that AI would be an important area for anyone studying computer science," Jørgen told us. But it wasn't just the topic that drew him in — it was the people. After looking up the founders and seeing their backgrounds from places like MIT, Harvard, and Berkeley, he knew this was the place to be. "Even though the founders were more experienced, they were genuinely interested in helping newer students learn and grow," he recalls. That supportive environment paid off quickly. As a first-year student, Jørgen was thrown into a real-world project with Kjeldsberg Eiendomsforvaltning. It wasn't a textbook exercise with a clear answer; it was a challenging time-series project involving everything from XGBoost to Variational Autoencoders (VAEs). "I learned a lot about what it is actually like to work on a real machine learning project, including the ambiguity and practical challenges that come with it," he says. This hands-on experience was crucial, eventually helping him land his first summer internship as a Machine Learning Engineer at Å Energi. ## The Berkeley experience: "night culture" and high intensity Moving from NTNU to UC Berkeley brought a shift in pace and culture. Jørgen describes the environment in the U.S. as incredibly inspiring, but significantly more intense. "The workload at Berkeley is significantly higher," he admits. Unlike the Norwegian model where grades often hinge on a single final exam, Berkeley demands consistent performance throughout the semester. But the biggest surprise? The "night culture." "The library can be almost empty around 08:00, but completely full late at night, sometimes as late as 03:00," Jørgen says. "This feels very different from Norway, where students often compete for study spots early in the morning." However, the academic quality makes up for the sleepless nights. Jørgen highlights the "first principles" teaching style of the professors and the unique "DeCals" — student-led courses where you can earn credit learning niche topics from peers. ![Lecture hall in Jørgen's COMPSCI 188 class at UC Berkeley](/blog/jorgen/lecture.webp) Caption: Image from Jørgen's COMPSCI 188 — Introduction to Artificial Intelligence class. ## A hub for research Beyond the shift in study hours, Jørgen points out another fundamental difference: the deep integration of research into the student experience. At Berkeley, engaging in research is far more common for both undergraduate and graduate students than it is back home. "In several courses, conducting a genuine research project and writing an associated paper is actually part of the assessment," he explains. Students are actively encouraged to get involved from the very start of their studies. This is made possible not only because there are far more research environments and laboratories at Berkeley than at NTNU, but also because these labs are much more willing to bring in non-PhD students as research assistants. Jørgen believes this relentless focus on student involvement is a key reason why the university is a global powerhouse in his field. "With this strong commitment and focus on research — both through teaching and lab opportunities — it is not surprising that Berkeley is among the world's leading universities in Machine Learning and Artificial Intelligence," Jørgen says. "This is something NTNU can definitely learn from, and which I believe Norway as a whole would greatly benefit from." ## Hacking in San Francisco You can't study near Silicon Valley without getting your hands dirty in the tech scene. Jørgen recently participated in the Hacktoberfest hackathon in San Francisco, where his team aimed to build something with real business value: automating fuzz testing using AI. Leveraging Large Language Models (LLMs), they built a solution that was technically feasible and commercially viable. Jørgen credits his time in ReLU for giving him the confidence to tackle the challenge. "My experience from ReLU helped me feel comfortable setting up an AI project, choosing appropriate models, and working with LLMs," he explains. ## Lessons for Norwegian students Reflecting on the differences between the two worlds, Jørgen believes students back home could learn a thing or two from the American mindset. "I think Norwegian students could benefit from being a bit more hungry, optimistic, and willing to believe that they can contribute, even without knowing everything in advance," he says. He also notes that the competitive nature of Berkeley tech clubs — some requiring four to five rounds of interviews — pushes students to be sharper and more driven. ## Looking forward So, where is AI going next? Jørgen sees a shift toward integration. "I strongly believe that AI will become more integrated with companies' own data… This is already happening, and Norway has several companies that are well positioned in this space." Jørgen's advice for those wanting to follow a similar path: - **Be genuinely interested.** You have to want to keep up with developments. - **Learn the math.** It's the foundation of everything in ML. - **Be independent.** Clubs like ReLU are great starters, but you need to put in the extra effort on your own. "Approach large and difficult projects with an open and confident mindset," Jørgen concludes. "Many students convince themselves that they are not good enough to contribute, and with that mindset, you learn very little." --- ### Announcing collaborations with NAIL and AID URL: /blog/nail-aid-collaboration/ Date: 2025-12-16 | Author: Board ReLU NTNU is partnering with the Norwegian Open AI Lab (NAIL) and the AI for Decisions (AID) center to bridge the gap between students, academia, and industry. ReLU NTNU is thrilled to announce collaborations with the Norwegian Open AI Lab (NAIL) and the AI for Decisions (AID) center. This marks a significant step forward in our mission to bridge the gap between students, academia, and industry, creating a powerhouse of AI innovation and talent development in Norway. At ReLU NTNU, we are dedicated to providing students with hands-on experience in solving real-world challenges with machine learning. By joining forces with NAIL and AID, we are positioned to create even more opportunities for students to engage with cutting-edge research, network with industry professionals, and develop their skills in this rapidly evolving field. ## Norwegian Open AI Lab ![Norwegian Open AI Lab logo](/blog/nail-aid/nail.webp) NAIL is a hub for research, education, and innovation in AI. The lab is hosted at NTNU in Trondheim and works as a network between academia, students, businesses, and public sector organizations. [Read more about NAIL](https://www.ntnu.edu/ailab) ## AI for Decisions ![AI for Decisions logo](/blog/nail-aid/aid.webp) AID is one of six new national research centers in artificial intelligence funded by the Norwegian government through the Research Council and its industry partners. The center is led by NTNU in collaboration with SINTEF. Under the direction of Professor Sebastien Gros (NTNU) and Signe Riemer-Sørensen (SINTEF), AID is dedicated to researching how AI can be leveraged to handle risk and improve decision-making in complex and critical situations. [Read more about AID](https://www.linkedin.com/company/aidcenter/) ## A shared vision These partnerships are more than just a formal connection; they represent a shared vision for the future of AI in Norway. Through its collaborations with NAIL and AID, ReLU NTNU will create a powerful synergy that benefits students, researchers, and industry partners alike. Our combined efforts will focus on: - **Enhancing student engagement** — creating new avenues for students to participate in AI-related events, workshops, and projects. - **Fostering collaboration** — facilitating networking opportunities between students, researchers, and industry leaders. - **Promoting knowledge sharing** — disseminating the latest advancements in AI and providing platforms for open discussion and learning. We are incredibly excited about the potential of these collaborations and look forward to working closely with NAIL and AID. Together, we will empower the next generation of AI leaders and solidify NTNU's position at the forefront of artificial intelligence. Stay tuned for more updates on the exciting initiatives that will emerge from these partnerships. --- ### Welcoming the 2025 cohort: kickoff weekend URL: /blog/2025-kickoff-weekend/ Date: 2025-09-09 | Author: Board ReLU NTNU officially launched the 2025 academic year by welcoming a new cohort of talented and driven students. Two days of presentations, workshops, team building, and an Estimathon. This past weekend, ReLU NTNU officially launched the 2025 academic year by welcoming a new cohort of talented and driven students. Our annual kickoff event, held on September 6th and 7th, was a dynamic and immersive start, setting the stage for a year of ambitious projects, intensive learning, and impactful collaboration. ## Day 1: Building the foundation Saturday was dedicated to integrating our new members into the heart of our organization. The day began with a series of presentations from the ReLU Board, who introduced the collective's core mission: to foster future leaders in artificial intelligence by applying machine learning to solve real-world industry challenges. New members received a comprehensive overview of our operations, from the organization's economic structure to a detailed walkthrough of our unique AI Educational Program. The sessions provided a clear roadmap of the skills and experience they will gain throughout the year. The afternoon shifted from strategy to the practical fundamentals required for success. After a fun, competitive quiz, the focus turned to essential tools and protocols. This included a hands-on workshop dedicated to setting up coding environments and a crucial briefing on navigating information systems and adhering to professional standards with NDAs. The day concluded with a team dinner, providing a perfect opportunity for new and existing members to forge connections in a relaxed, informal setting. ![ReLU Quiz 3 projected on screen with members seated at tables.](/blog/2025-kickoff/quiz.webp) ## Day 2: Fostering community and embracing the challenge Sunday was all about strengthening community bonds and diving into the hands-on work that defines us. The morning began with team-building activities and icebreakers, reinforcing the collaborative spirit that is crucial to our success. ![Members in a circle playing a name game during Sunday's icebreakers.](/blog/2025-kickoff/name_game.webp) The highlight of the day was an inspiring showcase of impactful projects from the previous year. This session gave our new cohort a tangible look at the complex challenges and innovative solutions ReLU teams deliver in partnership with leading companies. Building on that momentum, the weekend culminated in our first Estimathon of the year. This hands-on challenge saw our new members immediately put their problem-solving skills to the test in a fast-paced, collaborative environment. ![Two members presenting in front of an Estimathon problem set.](/blog/2025-kickoff/estimathon.webp) We are incredibly impressed by the enthusiasm and talent of our new members. The kickoff weekend was a resounding success, and our team is now fully assembled and ready to tackle the challenges that lie ahead. Welcome to ReLU NTNU. We are excited to see the incredible solutions you will build and the leaders you will become. --- ## Project Archive ### Fine-tuning open-source LLMs for energy news summarization Partner: Rystad Energy Semester: Spring 2024 Status: completed Areas: LLMs, Summarization Methods: LoRA, Llama 2, Llama 3, Gemma, GPT-4 benchmark Fine-tuned open-source LLMs with LoRA for energy news summarization. Benchmarking against closed-source models was how the team evaluated whether generated summaries matched GPT-4-level quality. Data: News articles Handoff: code; report --- ### Finding drivers of salmon appetite Partner: Grieg Seafood Semester: Spring 2024 Status: completed Areas: Time series, Explainability Methods: XGBoost, SHAP Built appetite models on feeding time series with XGBoost and SHAP, but the real goal was XAI: which features drive feeding behavior. Grieg already had models and prior insight—the project tested whether we could surface drivers they had not already identified. Data: Salmon farming feeding data Handoff: report; code --- ### Extracting people, countries, and projects from long-form articles Partner: Rystad Energy Semester: Spring 2024 Status: completed Areas: LLMs, Information extraction Methods: Llama 2, Llama 3, Gemma, GPT-4 Compared open and private LLMs for named entity extraction from long articles, with specialized metrics on extraction quality and a writeup of performance, quality, and economic viability. Data: Long-form articles Handoff: code; report --- ### Anomaly detection on building energy data, phase 1 Partner: Kjeldsberg Eiendomsforvaltning Semester: Spring 2024 Status: completed Areas: Time series, Anomaly detection Methods: generative models, data preprocessing First phase of a multi-semester anomaly detection project: data processing and model exploration on building energy consumption, including generative models. Data: Building energy consumption time series Handoff: report; code --- ### Entity matching across datasets Partner: Rystad Energy Semester: Fall 2024 Status: completed Areas: Entity matching, LLMs Methods: fuzzy search, embeddings, pairwise LLM matching, prompt engineering Tested preprocessing, fuzzy search, embeddings, pairwise LLM matching, and prompting strategies. Measured accuracy and Top-N accuracy and produced a report on performance, results, and economic viability. Data: Entity records Evaluation: Accuracy, Top-N accuracy Handoff: code; report --- ### Explaining harvest deviation in salmon farming Partner: Grieg Seafood Semester: Fall 2024 Status: completed Areas: Time series, Explainability Methods: XGBoost, SHAP, ElasticNet Studied harvest deviation with an emphasis on interpretability. Combined XGBoost and SHAP with ElasticNet, then handed over a report with findings and code. Data: Salmon farming time series Handoff: report; code --- ### Validating customs declarations with OCR and language models Partner: NORBIT Semester: Fall 2024 Status: exploratory Areas: Document processing, OCR, LLMs Methods: OCR tools, LLM text structuring, layout-aware region classification Explored whether OCR plus language models could support validation of customs declaration documents. Looked at OCR tools, LLM-based interpretation of extracted text, fine-tuning for document understanding, and layout-aware region classification. Data: Customs declaration documents --- ### Aorta segmentation from abdominal CT scans Partner: Oslo University Hospital Semester: Fall 2024 Status: completed Areas: Medical imaging, Segmentation Methods: MONAI, 3D Slicer, 2D U-Net, Auto3Dseg, TotalSegmentator Explored segmentation of the abdominal aorta from routine CT scans to support detection of aneurysms. Tested several model families and built tooling around preprocessing, visualization, and post-processing for measurement. Data: Abdominal CT scans Handoff: preprocessing tools; visualization tools; post-processing analysis Public limits: Clinical data is not shown publicly. Patient information was not used in this writeup. --- ### Image captions and alt text for VG, phase 1 Partner: Schibsted Semester: Fall 2024 Status: completed Areas: Multimodal, LLMs Methods: Gemma, Llama, Qwen, multimodal fine-tuning, GPT-4 Turbo benchmark Fine-tuned Gemma, Llama, and Qwen for VG-style captions, compared multimodal and text-only fine-tuning, benchmarked against GPT-4 Turbo, and explored prompting for alt-text generation. Data: Article text and images Public limits: Copyright, newsroom workflow, and face-recognition details are not described publicly. --- ### Anomaly detection on building energy data, phase 2 Partner: Kjeldsberg Eiendomsforvaltning Semester: Fall 2024 Status: completed Areas: Time series, Anomaly detection Methods: statistical models Continued the anomaly detection project. Selected a statistical model that performed well and was easier to interpret than the generative baselines from phase 1. Data: Building energy consumption time series Handoff: pipeline; model selection report --- ### LLM agents for data retrieval and analysis, phase 1 Partner: Solution Seeker Semester: Fall 2024 Status: completed Areas: Agents, LLMs Methods: Plotly Dash, OpenAI Agents API, MCP, tool use Started a chat-based data exploration tool: an LLM agent that could generate code and database queries, retrieve data, and produce visualizations or statistical analyses through Plotly Dash. Data: Production data via partner systems Handoff: proof-of-concept app; documentation --- ### Increasing entity density in LLM summaries Partner: Rystad Energy Semester: Fall 2024 Status: completed Areas: LLMs, Summarization Methods: DocETL, prompt engineering Tested prompting techniques for entity-dense summaries of news articles using DocETL, with notes on speed, economic viability, and suggestions for automated evaluation. Data: News articles Handoff: Python module; report --- ### Generalized evaluation harness for entity matching pipelines Partner: Rystad Energy Semester: Spring 2025 Status: completed Areas: Entity matching, Evaluation, LLMs Methods: Weights & Biases, single-step pipelines, multi-step pipelines, prompting strategies Continued the prior semester's entity-matching work. Built a system for comparing entity-matching pipelines across datasets of varying complexity, combining Weights & Biases with single-step and multi-step strategies. Data: Entity records across datasets Evaluation: Quantitative testing across datasets Handoff: evaluation system; report; usage instructions --- ### Classifying salmon harvest deviation Partner: Grieg Seafood Semester: Spring 2025 Status: completed Areas: Time series, Classification Methods: ARIMA, XGBoost, LSTM, ROCKET, ensemble Reframed harvest deviation as a classification problem: predicting whether end-of-cycle deviation would be high or low. Explored ARIMA, XGBoost, LSTM, and ROCKET, then handed over an ensemble model. Data: Salmon farming time series Handoff: ensemble model; report --- ### Sonar market signals from news Partner: NORBIT Semester: Spring 2025 Status: exploratory Areas: LLMs, Information extraction Methods: web scraping, GDELT, GPT models, Perplexity Sonar, few-shot prompting Experimental project on identifying sonar market drivers from news. Tested web scraping, GDELT, GPT models, Perplexity Sonar models, few-shot prompting, and reasoning-style prompting in a pipeline for collecting, filtering, and summarizing relevant articles. Data: News articles --- ### Image captions and alt text for VG, phase 2 Partner: Schibsted Semester: Spring 2025 Status: completed Areas: Multimodal, LLMs Methods: multimodal fine-tuning, prompt engineering Continued the captioning and alt-text work from Fall 2024 with Schibsted Media for VG. Extended fine-tuning experiments and prompting work for the newsroom use case. Data: Article text and images Public limits: Copyright, newsroom workflow, and face-recognition details are not described publicly. --- ### Anomaly detection on building energy data, phase 3 Partner: Kjeldsberg Eiendomsforvaltning Semester: Spring 2025 Status: completed Areas: Time series, Anomaly detection Methods: None listed Continued the Kjeldsberg anomaly-detection project, building on the interpretable statistical model selected in phase 2. Data: Building energy consumption time series --- ### LLM agents for data retrieval and analysis, phase 2 Partner: Solution Seeker Semester: Spring 2025 Status: completed Areas: Agents, LLMs Methods: None listed Continued the chat-based data exploration tool from phase 1: an LLM agent that retrieves data, runs analyses, and produces visualizations through partner systems. Data: Production data via partner systems --- ### Abdominal aortic aneurysm segmentation from CT scans Partner: Oslo University Hospital Semester: Spring 2025 Status: completed Areas: Medical imaging, Segmentation Methods: None listed Continued the aorta-segmentation work from Fall 2024, narrowing in on segmenting abdominal aortic aneurysms in routine CT scans. Data: Abdominal CT scans Public limits: Clinical data is not shown publicly. --- ### LLM-based extraction from Russian oil-and-gas reports Partner: Rystad Energy Semester: Fall 2025 Status: completed Areas: LLMs, Information extraction, Energy Methods: LLM-based extraction, prompt engineering Built a pipeline that uses LLMs to pull structured information out of Russian oil-and-gas reports and load it into Rystad Energy's database. Data: Russian oil-and-gas industry reports Handoff: extraction pipeline; report --- ### Hierarchical classification of customer documents Partner: Triplan Semester: Fall 2025 Status: completed Areas: Document classification, NLP, LLMs Methods: LLMs, classical machine learning, encoder fine-tuning Researched and tested approaches to simplify Triplan's document labelling workflow — extracting text from PDFs and comparing LLM, classical ML, and encoder fine-tuning options. Data: Customer documents --- ### Optimal B2B credit-limit policy and LLM-based fraud detection Partner: Two Semester: Fall 2025 Status: completed Areas: Anomaly detection, LLMs, Fintech Methods: credit optimization, LLM agents, risk modeling Two parallel tracks with the same partner — a data-driven credit-limit policy and an LLM-based fraud detection agent for real-time B2B payment risk. Data: B2B payment and transaction data --- ### Predicting maintenance needs for water pipelines Partner: Gemini Semester: Fall 2025 Status: completed Areas: Time series, Predictive maintenance Methods: survival analysis, failure classification, exploratory data analysis Exploratory phase of a predictive maintenance project for municipal water pipeline networks — structuring the problem and comparing survival analysis with failure classification. Data: Pipeline asset records and operational time series --- ### State of the art segmentation of AAA Partner: Oslo University Hospital Semester: Fall 2025 Status: completed Areas: Medical imaging, Segmentation, Clinical research Methods: segmentation baselines, preprocessing experiments Leveraging recent advances in medical image segmentation to identify potential abdominal aortic aneurysms from general abdominal CT scans acquired through the national screening program. Data: Abdominal CT scans Handoff: inference outputs; handoff notes for OUS review Public limits: Clinical data is not shown publicly. --- ### ESG data extraction using LLMs and RAG Partner: Rystad Energy Semester: Spring 2026 Status: completed Areas: LLMs, Information extraction, Energy, RAG Methods: LLM-based extraction, retrieval-augmented generation, audit logging Continued the Rystad collaboration with automated extraction of ESG data from oil and gas sustainability reports — comparing direct LLM extraction with a retrieval-augmented pipeline. Data: Oil and gas operator sustainability reports Handoff: full source code; metric-specific prompt templates; evaluation functions --- ### Hierarchical document classification, phase 2 Partner: Triplan Semester: Spring 2026 Status: completed Areas: Document classification, NLP, LLMs Methods: LLMs, classical machine learning, encoder fine-tuning Carried the fall research forward — further testing and refinement of the most promising classification approach for Triplan's large document labelling workflow. Data: Customer documents Handoff: refined classification system; evaluation results; integration notes --- ### Predictive maintenance for water pipelines, phase 2 Partner: Gemini Semester: Spring 2026 Status: completed Areas: Time series, Predictive maintenance Methods: XGBoost, feature engineering, spatial neighbor modeling Built an XGBoost-based pipeline on Stavanger municipality data to predict pipe failure risk across multiple time horizons, combining asset attributes with break history and spatial neighbor features. Data: Pipeline asset records and operational time series from Stavanger municipality Handoff: trained model; full modeling pipeline; documentation ---