במסגרת הכנס הארצי למדעי המחשב והמידע, אשר יתקיים ב־08/10/26, אנו נקיים מושב בבינה מלאכותית. המושב משקף את המיקוד של הכנס בתמורות בבינה מלאכותית באקדמיה, בתעשייה ובחברה בישראל. הוא יכלול הרצאות מוזמנות ממיטב החוקרים בתחום, ממגוון אוניברסיטאות ומוסדות מחקר. חלק מהמושב יוקדש למפגש השנתי של האיגוד הישראלי לבינה מלאכותית.
In this talk I will present how combining the power of Brains & Deep-Networks (DNNs) can lead to significant breakthroughs in both domains and potentially bridge the gap between Minds & Machines. I will further show how combining the power of Multiple Brains (“The Wisdom of a Crowd of Brains”) may lead to new breakthrough discoveries in Brain-Science, allow mapping of information between different brains (with NO shared data), and lead to new ways of training and interpreting artificial DNNs.
Novel Lossy Image Compression Schemes via Diffusion Models
Abstract
Diffusion models have transformed the landscape of image generation, offering a reliable sampling from the probability density function of images. These models have also been extended to provide posterior sampling capabilities, by conditioning the process on given external controls. This ability has been demonstrated for handling inverse problems in novel ways, and for producing images conditioned on textual prompts, showing in both tasks remarkable results. A very new and fascinating trend in this domain harnesses diffusion models to image compression, showing great potential here as well. This talk presents this story: The essence of diffusion models, their generative abilities, their posterior-sampling extensions, and most importantly, their ability to lead to novel and highly effective compression schemes for images and other data sources. No background is assumed.
Language models (LMs) capture large amounts of factual knowledge applicable to a wide range of tasks. However, we do not yet have a coherent theory for how this knowledge is stored and retrieved. A compelling analogy is that of knowledge bases, which have traditionally been used to represent facts in a structured manner. In this talk, I will present our recent results that shed light on mechanisms for information storage and retrieval in LLMs, and how these relate to a knowledge base view.
The Challenge and Reward of Fair Play in Narrative: A Computational Investigation
Abstract
Effective storytelling relies on a delicate balance between meeting the reader's prior expectations and introducing unexpected developments. In the domain of detective fiction, this tension is known as fair play, which includes the implicit agreement between the writer and the reader as to the possibility and challenge of identifying the culprit through the story's clues. I will present a probabilistic framework that aims to quantify this elusive notion. Our definitions present an inherent tension between the coherence of the story, which measures how much the resolution “makes sense” in explaining the clues, and the surprise it induces. Due to this tension, balancing these qualities is challenging. We operationalize the proposed framework, defining an approximate computable metric of “fair play”.
Our validation simulations include both real stories and LLM-generated ones (the latter allowing maximal experimental control). The results present similar trends to the ones predicted by the framework. I will conclude by discussing potential broader implications.
This is joint work with Eitan Wagner and Renana Keydar.
To Grok Grokking: Provable Grokking in Ridge Regression
Abstract
Grokking is the sudden onset of generalization long after a model has already overfit its training data. In this talk, I will present an end-to-end theoretical account of this phenomenon in classical ridge regression. For over-parameterized linear models trained by gradient descent with weight decay, we prove the full grokking trajectory: early overfitting, a prolonged period of poor generalization, and eventual convergence to arbitrarily small generalization error. Our analysis provides quantitative bounds on the generalization delay, and shows how it can be amplified or eliminated through principled hyperparameter tuning. I will also present experiments suggesting that these predictions extend beyond linear models to nonlinear neural networks. Overall, the results suggest that grokking is not an inherent failure mode of deep learning, but rather a consequence of specific training conditions, and thus does not require fundamental changes to the model architecture or learning algorithm to avoid.
14:15–16:45 מושב אחר הצהריים: יו"ר המושב: ד"ר איתי ספרן.
14:15–14:30 פגישה שנתית של העמותה הישראלית לבינה המלאכותית. יו"ר: פרופ' טל גרינשפון.
14:30–15:10 הרצאת keynote של פרופ' מיכל פלדמן.
15:10–15:30 פרופ' עמרי אבנד.
15:30–15:50 ד"ר משה אליסוף.
15:50–16:10 פרופ' אמיר גלוברזון.
16:10–16:45 Poster session עם כיבוד קל.
Artificial Intelligence Session
In collaboration with the Israeli Association for Artificial Intelligence (IAAI).
As part of the National Conference on Computer and Information Sciences, taking place on 08/10/26, we will hold an artificial intelligence session. This session reflects the conference’s focus on developments in artificial intelligence in Israeli academia, industry, and society. It will feature invited talks by leading researchers in the field from various universities and research institutions. Part of the session will be devoted to the annual meeting of the Israeli Association for Artificial Intelligence.
In this talk I will present how combining the power of Brains & Deep-Networks (DNNs) can lead to significant breakthroughs in both domains and potentially bridge the gap between Minds & Machines. I will further show how combining the power of Multiple Brains (“The Wisdom of a Crowd of Brains”) may lead to new breakthrough discoveries in Brain-Science, allow mapping of information between different brains (with NO shared data), and lead to new ways of training and interpreting artificial DNNs.
Novel Lossy Image Compression Schemes via Diffusion Models
Abstract
Diffusion models have transformed the landscape of image generation, offering a reliable sampling from the probability density function of images. These models have also been extended to provide posterior sampling capabilities, by conditioning the process on given external controls. This ability has been demonstrated for handling inverse problems in novel ways, and for producing images conditioned on textual prompts, showing in both tasks remarkable results. A very new and fascinating trend in this domain harnesses diffusion models to image compression, showing great potential here as well. This talk presents this story: The essence of diffusion models, their generative abilities, their posterior-sampling extensions, and most importantly, their ability to lead to novel and highly effective compression schemes for images and other data sources. No background is assumed.
Language models (LMs) capture large amounts of factual knowledge applicable to a wide range of tasks. However, we do not yet have a coherent theory for how this knowledge is stored and retrieved. A compelling analogy is that of knowledge bases, which have traditionally been used to represent facts in a structured manner. In this talk, I will present our recent results that shed light on mechanisms for information storage and retrieval in LLMs, and how these relate to a knowledge base view.
The Challenge and Reward of Fair Play in Narrative: A Computational Investigation
Abstract
Effective storytelling relies on a delicate balance between meeting the reader's prior expectations and introducing unexpected developments. In the domain of detective fiction, this tension is known as fair play, which includes the implicit agreement between the writer and the reader as to the possibility and challenge of identifying the culprit through the story's clues. I will present a probabilistic framework that aims to quantify this elusive notion. Our definitions present an inherent tension between the coherence of the story, which measures how much the resolution “makes sense” in explaining the clues, and the surprise it induces. Due to this tension, balancing these qualities is challenging. We operationalize the proposed framework, defining an approximate computable metric of “fair play”.
Our validation simulations include both real stories and LLM-generated ones (the latter allowing maximal experimental control). The results present similar trends to the ones predicted by the framework. I will conclude by discussing potential broader implications.
This is joint work with Eitan Wagner and Renana Keydar.
To Grok Grokking: Provable Grokking in Ridge Regression
Abstract
Grokking is the sudden onset of generalization long after a model has already overfit its training data. In this talk, I will present an end-to-end theoretical account of this phenomenon in classical ridge regression. For over-parameterized linear models trained by gradient descent with weight decay, we prove the full grokking trajectory: early overfitting, a prolonged period of poor generalization, and eventual convergence to arbitrarily small generalization error. Our analysis provides quantitative bounds on the generalization delay, and shows how it can be amplified or eliminated through principled hyperparameter tuning. I will also present experiments suggesting that these predictions extend beyond linear models to nonlinear neural networks. Overall, the results suggest that grokking is not an inherent failure mode of deep learning, but rather a consequence of specific training conditions, and thus does not require fundamental changes to the model architecture or learning algorithm to avoid.