Rodrigo da Silva Alves

Rodrigo da Silva Alves

Welcome! Vítejte! Bem-vindo!
I am an Assistant Professor at the Department of Applied Mathematics
Faculty of Information Technology / Czech Technical University in Prague (FIT/CTU).

About me

Rodrigo da Silva Alves

Currently, I am an Assistant Professor (Recombee Lab's member) at the Department of Applied Mathematics, Faculty of Information Technology / Czech Technical University in Prague (FIT/CTU). My research focuses on recommender systems (theory and applications) and, more recently, sports analytics. However, I am interested in machine learning in general, especially related to data mining and how artificial intelligence correlates to other areas of computer science. Previously, I completed a Ph.D. in Computer Science at the Machine Learning Group under the supervision of Prof. Marius Kloft at TU Kaiserslautern , RP, Germany. I hold a Bachelor's degree in Information Systems and a Master's in Computer Science from the Department of Computer Science at the Federal University of Minas Gerais under the supervision of Prof. Renato Assunção. I am also a vocational educational teacher certificated by the Häme University of Applied Sciences, Finland, and I was a Lecturer from the Department of Applied Social Sciences at CEFET-MG. During my career, I had the pleasure of collaborating with various research groups (in different countries), outstanding students, and technical teams.

My hometown is Belo Horizonte (you can also say Beagá), Minas Gerais, Brazil. I am a Cruzeiro Esporte Clube fan and a enthusiast of Brazilian (multi-)culture and music, particularly Bossa Nova. I believe in the need of building an inclusive, accessible, multicultural, and prejudice-free environment for research.

Resume

Professional Experience

Assistant Professor (Odborný asistent)

Feb 2022 - present

Department of Applied Mathematics
Czech Technical University in Prague
Prague, Czech Republic

Head of Research

Feb 2022 - present

Recombee Lab
Recombee
Prague, Czech Republic

Researcher (Wissenschaftlicher Mitarbeiter)

Jan 2019 - Jan 2022

Machine Learning Group
Technical University of Kaiserslautern
Kaiserslautern, Rhineland-Palatinate, Germany

Lecturer (Professor EBTT)

Apr 2014 - Dec 2021 (On Ph.D. leave from Feb 2018)

Department of Applied Social Sciences
Centro Federal de Educação Tecnológica de Minas Gerais
Belo Horizonte, Minas Gerais, Brazil

Education

Dr. Rer. Nat - Computer Science

Feb 2018 - Feb 2022

Department of Computer Science
Technical University of Kaiserslautern
Kaiserslautern, Rhineland-Palatinate, Germany

Thesis: Towards Comprehensive Cluster-induced Methods for Recommender Systems
Supervisor: Prof. Marius Kloft

Professional Development Program for Teachers

Apr 2016 - Dec 2016

Häme University of Applied Sciences
Hämeenlinna, Finland

Development Work: Education for the Future: Applying Student-centered Learning in Brazilian Vocational Education
Supervisors: Dr. Essi Ryymin and Dr. Irma Kunnari
20 ECTS / 540 hours

Master of Computer Science

Feb 2013 - May 2015

Department of Computer Science
Federal University of Minas Gerais
Belo Horizonte, Minas Gerais, Brazil

Thesis: Stochastic point process mixing model for inter-event times of Web services
My Master's thesis is composed in Portuguese. Nevertheless, you may read this paper, which is an outcome of this research.
Supervisor: Prof. Renato Assunção
Co-Supervisor: Prof. Pedro O.S. Vaz de Melo

Bachelor of Information Systems

Feb 2009 - Dec 2012

Department of Computer Science
Federal University of Minas Gerais
Belo Horizonte, Minas Gerais, Brazil

Award: Best Student Award

News & Updates

[04/2026] "Learning Minimally Rigid Graphs with High Realization Counts" accepted at IJCAI-ECAI 2026

Happy to share that our paper on learning minimally rigid graphs with high realization counts, joint work with Oleksandr Slyvka, Jan Rubeš, and Jan Legerský, was accepted at IJCAI-ECAI 2026, the 35th International Joint Conference on Artificial Intelligence, in Bremen, Germany. A preprint is available on arXiv; proceedings are not out yet. See the Research section for details.

[03/2026] Three papers accepted at ACM UMAP 2026

Pleased to share that three papers were accepted at the 34th ACM Conference on User Modeling, Adaptation and Personalization (UMAP 2026): "The Stars Align: Modeling User Rating Calibration with Sparse Semantic Review Features", "Language Embeddings Meet Shallow Autoencoders", and "Leveraging Artist Catalogs for Cold-Start Music Recommendation". See the Research section for details.

[02/2026] Promoted to Area Chair for KDD

Honored to have been promoted to Area Chair for the ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD).

[01/2026] "Efficient Learning of Sparse Representations from Interactions" accepted at The Web Conference (WWW) 2026

Joint work with Vojtěch Vančura, Martin Spišák, and Ladislav Peška on learning high-dimensional sparse embedding layers for scalable retrieval in recommender systems was accepted at the ACM Web Conference 2026. See the Research section for details.

[09/2025] "Generalization bounds for rank-sparse neural networks" accepted at NeurIPS 2025

Our paper on generalization bounds exploiting the approximate low-rank structure of neural network weight matrices was accepted at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025), joint work with Antoine Ledent and Yunwen Lei.

[09/2025] Industry Track (OC) Chair at ACM RecSys 2025

Served as Industry Track Organizing Committee Chair for the Nineteenth ACM Conference on Recommender Systems (RecSys 2025).

[07/2025] Three papers at ACM RecSys 2025

Happy to have contributed to three papers at the Nineteenth ACM Conference on Recommender Systems (RecSys 2025): "Recurrent Autoregressive Linear Model for Next-Basket Recommendation", "Probabilistic Modeling, Learnability and Uncertainty Estimation for Interaction Prediction in Movie Rating Datasets", and "The Future is Sparse: Embedding Compression for Scalable Retrieval in Recommender Systems".

[03/2025] "SCORE: A convolutional approach for football event forecasting" published in the International Journal of Forecasting

My sole-authored paper on a convolutional approach to forecasting football match events was published in the International Journal of Forecasting.

[05/2024] "Generalization Analysis of Deep Non-linear Matrix Completion" accepted at ICML 2024

Our paper providing generalization bounds for Schatten p quasi-norm constrained matrix completion, joint work with Antoine Ledent, was accepted at the Forty-first International Conference on Machine Learning (ICML 2024).

[02/2024] "Unraveling the dynamics of stable and curious audiences in web systems" accepted at The Web Conference (WWW) 2024

Joint work with Antoine Ledent, Renato Assunção, Pedro Vaz-de-Melo, and Marius Kloft on modeling stable versus bursty "curious" audiences in web systems was accepted at WWW 2024.

[07/2023] 23rd European Agent Systems Summer School

Happy to announce my active participation in the 23rd European Agent Systems Summer School, hosted at the Faculty of Information Technology, Czech Technical University in Prague. For more comprehensive information, please visit the event's official website here. Additionally, you can access the slides from my presentation, which are now available here.

I am pleased to announce the publication of my latest academic paper on IEEE Transactions of Neural Networks and Learning Systems. The paper focuses on improving recommender systems (RSs) by addressing the issue of unequal amounts of noise in observed ratings. We propose a nuclear-norm-based matrix factorization method that leverages side information to estimate the uncertainty associated with each rating. By using this uncertainty as a weighting factor in our optimization process, we can effectively handle potentially erroneous or noisy ratings. For more details, click here.

We studied inductive matrix completion (matrix completion with side information) under an i.i.d. subgaussian noise assumption at a low noise regime, with uniform sampling of the entries. The paper will appear in the AAAI 2023 proceeedings, but currently you can access here the arxiv version of the paper.

[01/2023] My website is online :)

Finally I managed to update my website. It is a working product version that will be updated time to time.

Teaching

Winter 26/27

[FIT-SM1] Machine Learning Seminar 1
[NIE-PML] Personalised Machine Learning

Summer 25/26

[NI-ADM] Data Mining Algorithms
[NI-AML] Advanced Machine Learning

Winter 25/26

[NIE-PML] Personalised Machine Learning
[BI-SZ] Knowledge Engineering Seminar

Summer 24/25

[NI-ADM] Data Mining Algorithms
[NI-AML] Advanced Machine Learning

Winter 24/25

[NIE-PML] Personalised Machine Learning
[BI-SZ] Knowledge Engineering Seminar

Summer 23/24

[NIE-ADM] Data Mining Algorithms
[NI-ADM] Data Mining Algorithms
[NI-AML] Advanced Machine Learning

Winter 23/24

[NIE-ML1] Machine Learning 1
[NIE-PML] Personalised Machine Learning
[BI-SZ] Knowledge Engineering Seminar

Summer 22/23

[NIE-ADM] Data Mining Algorithms
[NI-ADM] Data Mining Algorithms
[NI-AML] Advanced Machine Learning

Winter 22/23

[BIE-VDZ] Data Mining
[BI-SZ] Knowledge Engineering Seminar

Summer 21/22

[NI-ADM] Data Mining Algorithms

Thesis Supervision

Bachelor's and Master's theses I have supervised at FIT/CTU, listed by year. See the official, continuously updated list on the FIT/CTU faculty page.

2026

Integrating User Signals for Enhancing Triplet-Based Cognitive Modeling

Ketevani Dzebniauri — Master's Thesis

Natural Language Steering of Recommender Systems via User Embedding Manipulation

Jaroslav Hradil — Master's Thesis

Incorporating Item Similarity into Exposure-Based Recommendation Models of User Preference Evolution

Jozef Koleda — Master's Thesis

Permutation-Equivariant Models for In-Game Event Prediction in Football

Václav Tran — Master's Thesis

Item Identifiability from Large-Language-Models-Based Embeddings in Recommender Systems

Linda Beková — Master's Thesis

2025

Leveraging Large Language Models for Regionalized Recommender Systems

Adam Čapka — Bachelor's Thesis

A Large Language Models Framework for Football Event Prediction

Dmytro Borovko — Bachelor's Thesis

An Artificial Intelligence-Based System for Automatic Reflection Question Generation in Educational Settings

Ondřej Holub — Bachelor's Thesis

Segment-Based Recommendations

Patrik Malý — Master's Thesis

Investigating Scoring and Ordering in Multi-Stage Recommender Systems for Book Recommendations Using Large Language Models

Maksim Spiridonov — Master's Thesis

2024

Multitask Learning for Cognitive Sciences Triplet Analysis

Tsimafei Stambrouski — Bachelor's Thesis

Harnessing Spatial Context for Item Recommendation

Vendula Švastalová — Master's Thesis

Human Alignment of Natural Language Processing Models

Anastasiia Solomiia Hrytsyna — Master's Thesis

Graph-Based Fraud Detection in Recommender Systems

Daniel Bohuněk — Master's Thesis

2023

Machine Learning-Based Prediction of Football Match Statistics

Ondřej Herman — Bachelor's Thesis

Football outcomes prediction with tensor completion embeddings

Martin Kostrubanič — Master's Thesis

Research

I am always open to collaborating on fascinating and challenging projects. If you are a CTU student and are looking for a Bachelor's or Master's project or would like to have your first steps in research, I would be glad to have a meeting with you and discuss the possibility of mentoring you in some projects in my research area. For internal and external collaboration, please get in touch with me by email. Below, you can find my main research contributions. Here is my google scholar profile.

Abstract: User ratings are often treated as comparable across users, although identical scores may reflect different experiences. We study whether ratings can be viewed as user-specific discretizations of a shared semantic continuum derived from review text. Our method maps reviews into sparse semantic features with a sparse autoencoder and learns user-specific filters for each rating level. On Amazon Electronics, the learned embeddings align along a shared low-dimensional rating axis. Users differ mainly in how they anchor and partition this continuum, while preserving its overall ordinal structure. These findings support a semantic view of calibration beyond scalar bias correction.
ACM Conference on User Modeling, Adaptation and Personalization

Abstract: Shallow autoencoders are appealing recommenders due to their simplicity, scalability, and competitive retrieval quality, but they struggle in strict cold-start settings where new items have no interactions. We propose an inductive shallow autoencoder that leverages item side information (language embeddings) by fixing the decoder to item features and learning only an encoder in the same semantic space. To prevent trivial self-reconstruction without enforcing a hard zero diagonal, we introduce diagonal gating: a leave-one-item-out objective that blocks the self-copy shortcut only for the item being updated while retaining context from the rest of the user history. An alternating-style optimization trains the model. Experiments on three real-world benchmarks show consistent gains over strong cold-start baselines, including other shallow autoencoders, and support lightweight (cross-domain) semantic user modeling.
ACM Conference on User Modeling, Adaptation and Personalization Best Paper Runner-Up

Abstract: The item cold-start problem poses a fundamental challenge for music recommendation: newly added tracks lack the interaction history that collaborative filtering (CF) requires. Existing approaches often address this problem by learning mappings from content features such as audio, text, and metadata to the CF latent space. However, previous works either omit artist information or treat it as just another input modality, missing the fundamental hierarchy of artists and items. Since most new tracks come from artists with previous history available, we frame cold-start track recommendation as “semi-cold” by leveraging the rich collaborative signal that exists at the artist level. We show that artist-aware methods can more than double Recall and NDCG compared to content-only baselines, and propose ACARec, an attention-based architecture that generates CF embeddings for new tracks by attending over the artist's existing catalog. We show that our approach has notable advantages in predicting user preferences for new tracks, especially for new artist discovery and more accurate estimation of cold item popularity.
ACM Conference on User Modeling, Adaptation and Personalization Best Paper Runner-Up

Abstract: For minimally rigid graphs, the same edge-length data can admit multiple realizations (up to translations and rotations). Finding graphs with exceptionally many realizations is an extremal problem in rigidity theory, but exhaustive search quickly becomes infeasible due to the super-exponential growth of the number of candidate graphs and the high cost of realization-count evaluation. We propose a reinforcement-learning approach that constructs minimally rigid graphs via 0- and 1-extensions, also known as Henneberg moves. We optimize realization-count invariants using the Deep Cross-Entropy Method with a policy parameterized by a Graph Isomorphism Network encoder and a permutation-equivariant extension-level action head. Empirically, our method matches the known optima for planar realization counts and improves the best known bounds for spherical realization counts, yielding new record graphs.
International Joint Conference on Artificial Intelligence

Abstract: Behavioral patterns captured in embeddings learned from interaction data are pivotal across various stages of production recommender systems. However, in the initial retrieval stage, practitioners face an inherent tradeoff between embedding expressiveness and the scalability and latency of serving components, resulting in the need for representations that are both compact and expressive. To address this challenge, we propose a training strategy for learning high-dimensional sparse embedding layers in place of conventional dense ones, balancing efficiency, representational expressiveness, and interpretability. To demonstrate our approach, we modified the production-grade collaborative filtering autoencoder ELSA, achieving up to 10× reduction in embedding size with no loss of recommendation accuracy, and up to 100× reduction with only a 2.5% loss. Moreover, the active embedding dimensions reveal an interpretable inverted-index structure that segments items in a way directly aligned with the model's latent space, thereby enabling integration of segment-level recommendation functionality (e.g., 2D homepage layouts) within the candidate retrieval model itself.
ACM Web Conference

Abstract: Selecting appropriate physical models is a critical yet difficult step in many areas of computational science and engineering. In multiphase Computational Fluid Dynamics (CFD), practitioners must choose among numerous closure model combinations whose performance varies strongly across flow conditions. Sub-optimal choices can lead to inaccurate predictions, simulation failures, and wasted computational resources, making model selection a prime candidate for data-driven decision support. This work formulates closure model selection as a cold-start recommender system problem in a high-cost scientific domain. We propose a hybrid recommendation framework that combines (i) metadata-driven case similarity and (ii) collaborative inference via matrix completion. The approach enables case-specific model recommendations for entirely new CFD cases using their descriptive features, while leveraging historical simulation results from similar cases. The methodology is evaluated on 13,600 simulations across 136 validation cases and 100 model combinations. Results show that the proposed hybrid recommender consistently outperforms popularity-based and expert-designed reference models and reduces regret across the investigated sparsities.
arXiv 2026

Abstract: Designing good reflection questions is pedagogically important but time-consuming and unevenly supported across teachers. This paper introduces a reflection-in-reflection framework for automated generation of reflection questions with large language models (LLMs). Our approach coordinates two role-specialized agents, a Student-Teacher and a Teacher-Educator, that engage in a Socratic multi-turn dialogue to iteratively refine a single question given a teacher-specified topic, key concepts, student level, and optional instructional materials. We evaluate the framework in an authentic lower-secondary ICT setting, using GPT-4o-mini as the backbone model and a stronger GPT-4-class LLM as an external evaluator. Dynamic stopping combined with contextual information consistently outperforms fixed 5- or 10-step refinement, and our two-agent protocol produces questions that are judged substantially more relevant and deeper, and better overall, than a one-shot baseline using the same backbone model.
arXiv 2026

Abstract: Sparse autoencoders (SAEs) have recently emerged as pivotal tools for introspection into large language models. SAEs can uncover high-quality, interpretable features at different levels of granularity and enable targeted steering of the generation process by selectively activating specific neurons in their latent activations. Our paper is the first to apply this approach to collaborative filtering, aiming to extract similarly interpretable features from representations learned purely from interaction signals. In particular, we focus on a widely adopted class of collaborative autoencoders (CFAEs) and augment them by inserting an SAE between their encoder and decoder networks. We demonstrate that such representation is largely monosemantic and propose suitable mapping functions between semantic concepts and individual neurons. We also evaluate a simple yet effective method that utilizes this representation to steer the recommendations in a desired direction.
arXiv 2026

Abstract: Email outreach remains a cornerstone of modern marketing, enabling direct, timely communication. However, this strategy faces significant personalization challenges, since new campaigns typically lack historical interaction data and rich side information. In this work, we propose a framework that combines collaborative-filtering (CF) signals derived from a shallow autoencoder (SAE) with a Thompson Sampling-based multi-armed bandit to dynamically select small batches of recipients for each email template. We show SAEs help balance exploration and exploitation by quantifying recipient informativeness and confidence, enabling efficient personalization without model retraining during active learning. To facilitate reproducibility and future research, we release a large dataset of almost 15 million recipient-message interactions. Our experiments show that our method outperforms multiple baselines in retrieval metrics while retaining interpretable model components.
ACM International Conference on Information and Knowledge Management

Abstract: The rapidly expanding closure model catalogue for multiphase Computational Fluid Dynamics (CFD) presents an increasingly pressing challenge: the selection of optimal closure models for new simulation scenarios. Conventional approaches to this model selection dilemma rely heavily on extensive domain expertise, on expensive trial-and-error model evaluations, and ongoing research efforts to develop universally applicable model sets. This paper introduces an alternative, collaborative paradigm for model selection: by constructing validation databases storing the CFD performance of case–model interactions, we aim to preserve the CFD experience and embed it within a data-driven recommender system framework. We present a first-of-its-kind performance matrix that enables the novel application of recommender system algorithms to the closure model selection problem, showcasing the power of matrix completion methods to predict optimal closure models from extremely sparse performance data.
AI Thermal Fluids

Abstract: Next-basket recommendation aims to predict the (sets of) items that a user is most likely to purchase during their next visit, capturing both short-term sequential patterns and long-term user preferences. However, effectively modeling these dynamics remains a challenge for traditional methods, which often struggle with interpretability and computational efficiency. In this paper, we propose ReALM, a Recurrent Autoregressive Linear Model that explicitly captures temporal item-to-item dependencies across multiple time steps. By leveraging a recurrent loss function and a closed-form optimization solution, our approach offers both interpretability and scalability while maintaining competitive accuracy. Experimental results on real-world datasets demonstrate that ReALM outperforms several state-of-the-art baselines in both recommendation quality and efficiency.
ACM Conference on Recommender Systems

Abstract: In this paper, we examine the hypothesis that the interactions recorded in many Recommendation Systems datasets are distributed according to a low-rank distribution, i.e. a mixture of factorizable distributions. Surprisingly, we find that on several popular datasets, a simple non-negative matrix factorization method equals or outperforms more modern methods such as LightGCN, which indicates that the sampling distribution over interactions is indeed low-rank. Furthermore, we mathematically prove that low-rank distributions are learnable with a sparse number of observations, arguably providing some of the first generalization bounds for recommender systems in the implicit feedback setting. Finally, we propose the theoretically grounded concept of empirical expected recall as an uncertainty estimate for probabilistic models of the recommendation task, and demonstrate its success in a setting where user-wise abstentions are allowed.
ACM Conference on Recommender Systems

Abstract: Industry-scale recommender systems face a core challenge: representing entities with high cardinality, such as users or items, using dense embeddings that must be accessible during both training and inference. However, as embedding sizes grow, memory constraints make storage and access increasingly difficult. We describe a lightweight, learnable embedding compression technique that projects dense embeddings into a high-dimensional, sparsely activated space. Designed for retrieval tasks, our method reduces memory requirements while preserving retrieval performance, enabling scalable deployment under strict resource constraints. Our results demonstrate that leveraging sparsity is a promising approach for improving the efficiency of large-scale recommenders.
ACM Conference on Recommender Systems

Abstract: We introduce a new convolutional AutoEncoder architecture for user modelling and recommendation tasks with several improvements over the state of the art. Firstly, our model has the flexibility to learn a set of associations and combinations between different interaction types in a way that carries over to each user and item. Secondly, our model is able to learn jointly from both the explicit ratings and the implicit information in the sampling pattern. It can also make separate predictions for the probability of consuming content and the likelihood of granting it a high rating if observed. Finally, we provide several generalization bounds for our model, among the first for auto-encoders in a Recommender Systems setting. In experiments on several real-life datasets, we achieve state-of-the-art performance on both the implicit and explicit feedback prediction tasks despite relying on a single model for both.
IEEE Transactions on Neural Networks and Learning Systems

Abstract: We propose a large language model explainability technique for obtaining faithful natural language explanations by grounding the explanations in a reasoning process. When converted to a sequence of tokens, the outputs of the reasoning process can become part of the model context and later be decoded to natural language as the model produces either the final answer or the explanation. To improve the faithfulness of the explanations, we propose to use a joint predict-explain approach, in which the answers and explanations are inferred directly from the reasoning sequence, without the explanations being dependent on the answers and vice versa. We demonstrate the plausibility of the proposed technique by achieving a high alignment between answers and explanations in several problem domains, and show that the proposed use of reasoning can also improve the quality of the answers.
World Conference on Explainable Artificial Intelligence

Abstract: The triplet-based odd-one-out problem, which involves trials where human subjects are asked to select the most different concept among three, is a well-studied task in cognitive sciences. With the release of a large triplet-based dataset, THINGS, there has been a recent surge in the popularity of machine learning models aimed at learning mathematical representations of object concepts, such as SPoSE, VICE, and CARE. The first two models learn representations by maximizing the similarity between the two most similar objects, while the latter diverges by directly learning the odd-one-out, making its embedding more distant. No prior attempts have integrated both paradigms, which are important for understanding object representation in cognitive science. In this paper, we propose MASTER, a multitask learning method for the triplet problem that encapsulates both paradigms. Our results demonstrate that our method not only better predicts the odd-one-out object but also provides insightful representations for studying these concepts. Furthermore, we studied the conditions under which each model performs better, offering valuable insights for future research on how these paradigms affect human understanding of object concepts.
Expert Systems with Applications

Abstract: Football (also known as soccer or association football) is the most popular sport in the world. It is a blend of skill and luck, making it highly unpredictable. To address this unpredictability, there has been a surge in popularity over the past decade in employing machine learning techniques for forecasting football-related features. Despite this progress, the existing body of work remains in its early stages, lacking the depth required to capture the intricate nuances of the sport. In this study, we introduce a convolutional approach designed to predict the occurrence of the next event in a football match, such as a goal or a corner kick, relying solely on easy-to-access past events for predictions. Our methodology adopts an online approach, meaning predictions can be computed during a live match. To validate our approach, we conduct a comprehensive evaluation against five baseline models, utilizing data from various elite European football leagues.
International Journal of Forecasting

Abstract: It has been recently observed in much of the literature that neural networks exhibit a bottleneck rank property: for larger depths, the activation and weights of neural networks trained with gradient-based methods tend to be of approximately low rank. In this paper we investigate the implications of this phenomenon for generalization. More specifically, we prove generalization bounds for neural networks which exploit the approximate low rank structure of the weight matrices if present. The final results rely on the Schatten p quasi norms of the weight matrices: for small p, the bounds exhibit a sample complexity that scales with the rank of the weight matrices rather than the raw parameter count. We also demonstrate in experiments that this bound outperforms both classic parameter counting and norm based bounds in the typical overparametrized regime.
Advances in Neural Information Processing Systems

Abstract: We present a segment-aware analytics pipeline designed to support real-time editorial decision-making in digital media platforms. The core of our method combines large language model (LLM) embeddings with sparse autoencoders to extract interpretable, up-to-date segments from news articles. These segments are continuously refreshed and integrated into the recommendation platform, providing the foundation for analytics dashboards aligned with editorial needs. This demo paper describes our experience deploying the pipeline at The Telegraph and illustrates how advanced representation learning can bridge recommendation systems and editorial workflows in fast-paced news environments.
INRA 2025

Abstract: Linear shallow autoencoders have gained popularity in collaborative filtering benchmarks for recommender systems due to their simplicity and strong performance. However, models like EASE struggle to scale to datasets with a large number of items. This paper addresses this limitation by evaluating scalable variants of shallow linear autoencoders, namely, the recently proposed ELSA and SANSA, on large-scale datasets representative of real-world applications. We further propose improvements to enhance their scalability on large, sparse datasets. Our evaluation goes beyond standard offline and online performance metrics, incorporating key industrial considerations such as training time and memory usage. By analyzing the effects of data sparsity and catalog size, we provide practical guidance to select the most suitable model for a given dataset.
ACM Transactions on Recommender Systems

Abstract: Large language models (LLMs) are sophisticated artificial intelligence systems designed to process and understand natural language at a complex level. This research addresses the growing interest in understanding the mechanisms of LLMs and evaluating their alignment with human cognition. We introduce an innovative alignment assessment strategy, utilizing an odd-one-out triplet-based task to investigate LLMs' representations against human object concept mental organization. Our methodology, which incorporates image captioning zero/few-shot learning accuracy scoring, is designed to evaluate language models' ability to predict similarities and differences between object concepts. A comprehensive experimental evaluation was conducted, involving four image captioning strategies, twenty-four LLMs across eight model families, and three scoring procedures. Our study explores the impact of description comprehensiveness on model-human representation alignment and analyzes how LLMs represent different levels of human judgment patterns.
ACM Transactions on Intelligent Systems and Technology

Abstract: Regionalization, also known as spatially constrained clustering, is an unsupervised machine learning technique used to identify and define spatially contiguous regions. In this work, we introduce a methodology to regionalize recommendation systems (RSs) based on a collaborative filtering approach. Two main challenges arise when performing regionalization on users' preferences in RSs: (1) unstructured data, as interactions are often scarce and observed at a smaller scale; and (2) the difficulty of evaluating the quality of the results. To address these challenges, our methodology relies on inductive matrix completion (IMC) to recover unknown entries of a rating matrix while utilizing region information to extract a region-based feature matrix. This enables us to derive more accurate recommendations that consider regionalized effects and discover interesting patterns in localized user behavior. We present a real-world case study illustrating the interpretable information our model can derive in terms of regionalized recommendation relevance.
ACM Transactions on Spatial Algorithms and Systems

Abstract: We propose the Burst-Induced Poisson Process (BPoP), a model designed to analyze time series data such as feeds or search queries. BPoP can distinguish between the slowly-varying regular activity of a stable audience and the bursty curious audience, often seen in viral threads. Our model consists of two hidden, interacting processes: a self-feeding process (SFP) that generates bursty behavior related to viral threads, and a non-homogeneous Poisson process (NHPP) with step function intensity that is influenced by bursts from the SFP. The NHPP models the normal background behavior, driven solely by the overall popularity of the topic among a stable audience. Through extensive empirical work, we have demonstrated that our model fits and characterizes a large number of real datasets more effectively than state-of-the-art models. Most importantly, BPoP can quantify the stable media channels over time, serving as a valuable indicator of their popularity.
ACM Web Conference

Abstract: In areas of machine learning such as cognitive modeling or recommendation, user feedback is usually context-dependent. In this article, we consider a classification task where each input consists of three items (a triplet), and predict which will be selected. Our aim is not only to return accurate predictions for the selection task, but also to additionally provide interpretable feature representations for context and individual items. To achieve this, we introduce CARE, a specialized neural network architecture that yields Context-Aware REpresentations based on observations of triplets alone. We demonstrate that, in addition to achieving state-of-the-art performance at the selection task, our model can produce meaningful feature representations for items, as well as context, using only triplet responses. In addition, we prove parameter counting and generalization bounds for our i.i.d. setting, demonstrating that apparent sample sparsity arising from the combinatorially large number of possible triplets is no obstacle to efficient learning.
IEEE Transactions on Neural Networks and Learning Systems

Abstract: Relevance-based ranking is a popular ingredient in recommenders, but it frequently struggles to meet fairness criteria because social and cultural norms may favor some item groups over others. A fair ranking should balance the exposure of items from advantaged and disadvantaged groups. To this end, we propose a novel post-processing framework to produce fair, exposure-aware recommendations. Our approach is based on an integer linear programming model maximizing the expected utility while satisfying a minimum exposure constraint. The model has fewer variables than previous work and thus can be deployed to larger datasets and allows the organization to define a minimum level of exposure for groups of items. We conduct an extensive empirical evaluation indicating that our new framework can increase the exposure of items from disadvantaged groups at a small cost of recommendation accuracy.
Expert Systems with Applications

Abstract: We provide generalization bounds for matrix completion with Schatten p quasi-norm constraints, which is equivalent to deep matrix factorization with Frobenius constraints. In the uniform sampling regime, the sample complexity scales like Õ(rn) where n is the size of the matrix and r is a constraint of the same order as the ground truth rank in the isotropic case. We then present a non-linear model, Functionally Rescaled Matrix Completion (FRMC), which applies a single trainable function to each entry of a latent matrix, and prove that this adds only negligible terms to the overall sample complexity, whilst experiments demonstrate that this simple model improvement already leads to significant gains on real data. We also provide extensions of our results to various neural architectures, thereby providing the first comprehensive uniform convergence PAC analysis of neural network matrix completion.
International Conference on Machine Learning

Abstract: We propose a robust recommender systems model which performs matrix completion and a ratings-wise uncertainty estimation jointly. Whilst the prediction module is purely based on an implicit low-rank assumption imposed via nuclear norm regularization, our loss function is augmented by an uncertainty estimation module that learns an anomaly score for each individual rating via a Graph Neural Network: data points deemed more anomalous are down-weighted when training the low-rank module. The whole model is trained end-to-end, allowing anomaly detection to draw on the supervised information available in the form of ratings. Our model's predictors thus enjoy the favourable generalization properties that come with being chosen from a small function space (i.e., low-rank matrices), whilst exhibiting the robustness to outliers that comes with deep learning methods. Experiments on various real-life datasets demonstrate that our model outperforms standard matrix completion and other baselines.
ACM Conference on Recommender Systems

Abstract: The evaluation of recommendation systems is a complex task. The offline and online evaluation metrics for recommender systems are ambiguous in their true objectives. The majority of recently published papers benchmark their methods using ill-posed offline evaluation methodology that often fails to predict true online performance. Because of this, the impact that academic research has on the industry is reduced. The aim of our research is to investigate and compare the online performance of offline evaluation metrics. We show that penalizing popular items and considering the time of transactions during the evaluation significantly improves our ability to choose the best recommendation model for a live recommender system. Our results, averaged over five large-size real-world live data procured from recommenders, aim to help the academic community to understand better offline evaluation and optimization criteria that are more relevant for real applications of recommender systems.
arXiv 2023 Best Paper Award

Abstract: In a recommender systems (RSs) dataset, observed ratings are subject to unequal amounts of noise. Some users might be consistently more conscientious in choosing the ratings they provide for the content they consume. Some items may be very divisive and elicit highly noisy reviews. In this article, we perform a nuclear-norm-based matrix factorization method which relies on side information in the form of an estimate of the uncertainty of each rating. A rating with a higher uncertainty is considered more likely to be erroneous or subject to large amounts of noise, and therefore more likely to mislead the model. Our uncertainty estimate is used as a weighting factor in the loss we optimize. To maintain the favorable scaling and theoretical guarantees coming with nuclear norm regularization even in this weighted context, we introduce an adjusted version of the trace norm regularizer which takes the weights into account. This regularization strategy is inspired from the weighted trace norm which was introduced to tackle nonuniform sampling regimes in matrix completion. Our method exhibits state-of-the-art performance on both synthetic and real life datasets in terms of various performance measures, confirming that we have successfully used the auxiliary information extracted.
IEEE Transactions on Neural Networks and Learning Systems

Abstract: We study inductive matrix completion (matrix completion with side information) under an i.i.d. subgaussian noise assumption at a low noise regime, with uniform sampling of the entries. We obtain for the first time generalization bounds with the following three properties: (1) they scale like the standard deviation of the noise and in particular approach zero in the exact recovery case; (2) even in the presence of noise, they converge to zero when the sample size approaches infinity; and (3) for a fixed dimension of the side information, they only have a logarithmic dependence on the size of the matrix. Differently from many works in approximate recovery, we present results both for bounded Lipschitz losses and for the absolute loss, with the latter relying on Talagrand-type inequalities. The proofs create a bridge between two approaches to the theoretical analysis of matrix completion, since they consist in a combination of techniques from both the exact recovery literature and the approximate recovery literature.
AAAI Conference on Artificial Intelligence

Abstract: Recently, the RS research community has witnessed a surge in popularity for shallow autoencoder-based CF methods. Due to its straightforward implementation and high accuracy on item retrieval metrics, EASE is potentially the most prominent of these models. Despite its accuracy and simplicity, EASE cannot be employed in some real-world recommender system applications due to its inability to scale to huge interaction matrices. In this paper, we proposed ELSA, a scalable shallow autoencoder method for implicit feedback recommenders. ELSA is a scalable autoencoder in which the hidden layer is factorizable into a low-rank plus sparse structure, thereby drastically lowering memory consumption and computation time. We conducted a comprehensive offline experimental section that combined synthetic and several real-world datasets. We also validated our strategy in an online setting by comparing ELSA to baselines in a live recommender system using an A/B test. Experiments demonstrate that ELSA is scalable and has competitive performance. Finally, we demonstrate the explainability of ELSA by illustrating the recovered latent space.
ACM Conference on Recommender Systems

Abstract: In this paper, we bridge the gap between the state-of-the-art theoretical results for matrix completion with the nuclear norm and their equivalent in inductive matrix completion: (1) In the distribution-free setting, we prove bounds improving the previously best scaling of $O(rd^2)$ to $\widetilde{O}(d^{3/2}\sqrt{r})$, where $d$ is the dimension of the side information and $r$ is the rank. (2) We introduce the (smoothed) adjusted trace-norm minimization strategy, an inductive analogue of the weighted trace norm, for which we show guarantees of the order $\widetilde{O}(dr)$ under arbitrary sampling. In the inductive case, a similar rate was previously achieved only under uniform sampling and for exact recovery. Both our results align with the state of the art in the particular case of standard (non-inductive) matrix completion, where they are known to be tight up to log terms. Experiments further confirm that our strategy outperforms standard inductive matrix completion on various synthetic datasets and real problems, justifying its place as an important tool in the arsenal of methods for matrix completion using side information.
Advances in Neural Information Processing Systems

Abstract: A reasonable assumption in recommender systems is that the rows (users) and columns (items) of the rating matrix can be split into groups (communities) with the following property: each entry of the matrix is the sum of components corresponding to community behavior and a purely low-rank component corresponding to individual behavior. We investigate (1) whether such a structure is present in real-world datasets, (2) whether the knowledge of the existence of such structure alone can improve performance, without explicit information about the community memberships. To these ends, we formulate a joint optimization problem over all (completed matrix, set of communities) pairs based on a nuclear-norm regularizer which jointly encourages both low-rank solutions and the recovery of relevant communities. Since our optimization problem is non-convex and of combinatorial complexity, we propose a heuristic algorithm to solve it. Our algorithm alternatingly refines the user and item communities through a clustering step jointly supervised by nuclear-norm regularization. The algorithm is guaranteed to converge. We performed synthetic and real data experiments to confirm our hypothesis and evaluate the efficacy of our method at recovering the relevant communities. The results shows that our method is capable of retrieving such an underlying (community behaviour + continuous low-rank) structure with high accuracy if it is present.
PMLR: NeurIPS Workshop on Pre-registration in Machine Learning

Abstract: In this paper, we introduce a non-stationary and context-free Multi-Armed Bandit (MAB) problem and a novel algorithm (which we refer to as BMAB) to solve it. The problem is context-free in the sense that no side information about users or items is needed. We work in a continuous-time setting where each timestamp corresponds to a visit by a user and a corresponding decision regarding recommendation. The main novelty is that we model the reward distribution as a consequence of variations in the intensity of the activity, and thereby we assist the exploration/exploitation dilemma by exploring the temporal dynamics of the audience. To achieve this, we assume that the recommendation procedure can be split into two different states: the loyal and the curious state. We identify the current state by modelling the events as a mixture of two Poisson processes, one for each of the possible states. We further assume that the loyal audience is associated with a single stationary reward distribution, but each bursty period comes with its own reward distribution. We test our algorithm and compare it to several baselines in two strands of experiments: synthetic data simulations and real-world datasets. The results demonstrate that BMAB achieves competitive results when compared to state-of-the-art methods.
ACM Conference on Recommender Systems

Abstract: We propose orthogonal inductive matrix completion (OMIC), an interpretable approach to matrix completion based on a sum of multiple orthonormal side information terms, together with nuclear-norm regularization. The approach allows us to inject prior knowledge about the singular vectors of the ground-truth matrix. We optimize the approach by a provably converging algorithm, which optimizes all components of the model simultaneously. We study the generalization capabilities of our method in both the distribution-free setting and in the case where the sampling distribution admits uniform marginals, yielding learning guarantees that improve with the quality of the injected knowledge in both cases. As particular cases of our framework, we present models that can incorporate user and item biases or community information in a joint and additive fashion. We analyze the performance of OMIC on several synthetic and real datasets. On synthetic datasets with a sliding scale of user bias relevance, we show that OMIC better adapts to different regimes than other methods. On real-life datasets containing user/items recommendations and relevant side information, we find that OMIC surpasses the state of the art, with the added benefit of greater interpretability.
IEEE Transactions on Neural Networks and Learning Systems

Abstract: Activity coefficients, which are a measure of the nonideality of liquid mixtures, are a key property in chemical engineering with relevance to modeling chemical and phase equilibria as well as transport processes. Although experimental data on thousands of binary mixtures are available, prediction methods are needed to calculate the activity coefficients in many relevant mixtures that have not been explored to date. In this report, we propose a probabilistic matrix factorization model for predicting the activity coefficients in arbitrary binary mixtures. Although no physical descriptors for the considered components were used, our method outperforms the state-of-the-art method that has been refined over three decades while requiring much less training effort. This opens perspectives to novel methods for predicting physicochemical properties of binary mixtures with the potential to revolutionize modeling and simulation in chemical engineering.
The Journal of Physical Chemistry Letters

Abstract: In the so-called Total Quality Era, it is necessary to implement standardized and recognized experimental procedures around the world. When testing laboratories are adapted to the requirements set forth in ISO/IEC 17025:2017 standard, the evaluation of results and the exchange of knowledge becomes easier and more dynamic. This adaptation can be simplified and accelerated through the use of a data management software. Thus, the objective of this work was to develop a platform for quality control of a chemical testing laboratory, focusing on compliance with managerial and technical requirements of ISO/IEC 17025:2017 standard. The developed software allows not only data recording, but also the comparison of the analysis results with limit values established by current legislation, guaranteeing greater reliability of the reports issued. The created prototype is useful in ensuring high efficiency of the activities of chemical testing laboratories, making the workflow faster and safer, aside from guaranteeing compliance with the requirements of ISO/IEC 17025:2017 standard.
Brazilian Journal of Analytical Chemistry

Abstract: The problem to accurately and parsimoniously characterize random series of events (RSEs) seen in the Web, such as Yelp reviews or Twitter hashtags, is not trivial. Reports found in the literature reveal two apparent conflicting visions of how RSEs should be modeled. From one side, the Poissonian processes, of which consecutive events follow each other at a relatively regular time and should not be correlated. On the other side, the self-exciting processes, which are able to generate bursts of correlated events. The existence of many and sometimes conflicting approaches to model RSEs is a consequence of the unpredictability of the aggregated dynamics of our individual and routine activities, which sometimes show simple patterns, but sometimes results in irregular rising and falling trends. In this paper we propose a parsimonious way to characterize general RSEs, namely the Burstiness Scale (BuSca) model. BuSca views each RSE as a mix of two independent process: a Poissonian and a self-exciting one. Here we describe a fast method to extract the two parameters of BuSca that, together, gives the burstiness scale ψ, which represents how much of the RSE is due to bursty and viral effects. We validated our method in eight diverse and large datasets containing real random series of events seen in Twitter, Yelp, e-mail conversations, Digg, and online forums. Results showed that, even using only two parameters, BuSca is able to accurately describe RSEs seen in these diverse systems, what can leverage many applications.
ACM International Conference on Knowledge Discovery and Data Mining

Abstract: With the advancement of information systems, means of communications are becoming cheaper, faster, and more available. Today, millions of people carrying smartphones or tablets are able to communicate practically any time and anywhere they want. They can access their e-mails, comment on weblogs, watch and post videos and photos (as well as comment on them), and make phone calls or text messages almost ubiquitously. Given this scenario, in this article, we tackle a fundamental aspect of this new era of communication: How the time intervals between communication events behave for different technologies and means of communications. Are there universal patterns for the Inter-Event Time Distribution (IED)? How do inter-event times behave differently among particular technologies? To answer these questions, we analyzed eight different datasets from real and modern communication data and found four well-defined patterns seen in all the eight datasets. Moreover, we propose the use of the Self-Feeding Process (SFP) to generate inter-event times between communications. The SFP is an extremely parsimonious point process that requires at most two parameters and is able to generate inter-event times with all the universal properties we observed in the data. We also show three potential applications of the SFP: as a framework to generate a synthetic dataset containing realistic communication events of any one of the analyzed means of communications, as a technique to detect anomalies, and as a building block for more specific models that aim to encompass the particularities seen in each of the analyzed systems.
ACM Transactions on Knowledge Discovery from Data

Contact

Location:

Room A-1354 / Building A, 13th floor
Thákurova 7
Prague 6 – Dejvice
160 00

Please do not hesitate to contact me. I am often in my office, and you can visit me without an appointment. However, I am also frequently busy, so if you want to make sure you can talk to me, send a message before. If you are a CTU student looking for projects, read about it here.