Praveen explains why his team moved off GraphRAG in favor of a tree-structured memory for faster incremental updates and how they balance retrieval freshness, latency budgets, and access control at LinkedIn’s scale.
Connect with Praveen on LinkedIn.
Ryan Donovan (00:00)
Hello, everyone, and welcome to the Stack Overflow Podcast, a place to talk all things software and technology. I’m your host, Ryan Donovan, and today we are talking about a very large agentic memory operation that LinkedIn put together. My guest is Praveen Bodigutla, principal AI researcher at LinkedIn. Welcome to the show, Praveen.
Praveen Bodigutla (00:35)
Thank you, very nice to meet you, Ryan.
Ryan Donovan (00:37)
Glad you could be here. Before we get into the topic today, tell us a little bit about how you got into software and technology.
Praveen Bodigutla (00:47)
I’m a principal AI researcher at LinkedIn and I lead a few foundation efforts, including the memory agent, as well as serving as the founding engineer and leading some of the AI product development for both the enterprise and the consumer side. The journey to where I am right now is pretty nonlinear. I started as a platform engineer at Yahoo after finishing my undergrad in computer science and economics. Then I did a couple of master’s degrees — in financial mathematics and data science — from Stanford and NYU, where I got the opportunity to work with renowned professors such as Andrew Ng, Kyun Kyoon Chou, and Jeffrey Ullman. In my past life, I was a quant developer at an investment bank in New York, and after that I joined Alexa, where I worked on dialogue model research. The central theme of my career has been innovating on the AI side and building platforms so that innovation translates into useful products for end users.
Ryan Donovan (01:58)
We’re talking about something today that feels like a big innovation in a space a lot of people are discussing — agentic memory and context layers. You all built a whole cognitive memory agent at LinkedIn. Can you give us an overview of what the project was?
Praveen Bodigutla (02:25)
Let’s start with why we actually built the cognitive memory agent. LinkedIn developed and successfully launched a hiring agent — a hiring assistant for recruiters to manage their hiring workflows. Through those interactions, we observed that recruiters expressed hiring preferences and refined the roles they were hiring for. They mentioned where they were hiring, what skills they were interested in, and gave direct feedback on candidates shown to them. We noticed there was stickiness to these preferences, which translated into how they defined roles across similar titles and other positions. To provide a truly agentic experience for recruiters, we wanted to provide a personalization layer — and that is the genesis of the cognitive memory agent. It gives our agents a concept of state for the users, in this case recruiters, who are interacting with them.
We didn’t just want to fetch context; we wanted to manage the entire lifecycle of memory — the whole memory flywheel — starting from understanding what’s ingested, what’s retrieved, how it is contextually relevant, and how it is organized and updated. The memory agent manages this entire flywheel and provides deep personalization for our end users.
Ryan Donovan (04:13)
It’s interesting you mentioned state, and it’s not just conversational session state or preferences state — it’s a whole layered structure. Can you talk about what those layers are and why you needed them?
Praveen Bodigutla (04:37)
The memory layers we use within our memory agent are four in total. First is conversation memory — the most recent, up-to-date information from the current interaction, capturing the tasks and preferences being expressed in real time.
On the other end is the semantic memory layer, which is the aggregated information about a user based on interactions and preferences expressed across sessions. On LinkedIn, we have different product offerings and surfaces where users interact, and recruiters not only use the hiring assistant agent but also use a search platform to look for candidates. The semantic layer aggregates information not just from agent interactions but also from related activities across product services.
In between, we have the episodic store, which provides a temporal querying layer where we can identify the most recent relevant activities a user performed. It provides specificity of signal and also provenance — if we aggregate information, we can trace it back to what activities led us to conclude a particular preference.
Finally, there is the procedural memory layer. Even if recruiters are hiring for similar roles, the way they interact, the trade-offs they make, and the preferences they express are very different. Some may weight location and workplace type heavily, while others focus more on seniority or specific skills. The procedural memory captures the how of how the user achieves their tasks. Together, these form a layered cake of conversational, episodic, procedural, and semantic memory that captures information at different levels of granularity and specificity.
Ryan Donovan (07:09)
All of this comes from essentially a single data stream — the discussions and interactions on the site. What’s the challenge of extracting these separate behavioral state markers?
Praveen Bodigutla (07:38)
There are two different sets of data sources. One is the interaction the recruiter or user has with the agents. As you accumulate interactions over time, your context can bloat. Think about the hiring process involving multiple steps: starting with the job description, refining it, reviewing candidates, giving feedback, reaching out to some and archiving others. All of this accumulates. At runtime, we need to make sure this is compacted correctly.
We have an ingestion service that looks at interactions in real time and near-real time and tries to consolidate them. In fixed workflow agents, this is relatively easier because you know the boundaries of each step. But as we move toward more advanced agents, the steps can be interlinked — a recruiter might start by calibrating a candidate, then go back and refine the job description, then loop between the two. Identifying the right session boundaries, identifying the right interaction subtopics, organizing the memory, and making sure it is discoverable — that stream is challenging in its own right. You have to compress it without losing information and accurately represent what is persisted.
Then there is retrieval. As an interaction is ongoing, preferences can change — a recruiter might say they no longer want to hire in a particular location, or they want to look for a complementary skill set. When preferences shift mid-conversation, we have to ensure the most fresh and recent information gets priority. If there are conflicts while retrieving memory and answering queries posed through the application agent, those conflicts need to be handled correctly, with accurate and fresh information, and in low latency.
Ryan Donovan (10:07)
A recruiter could be hiring for multiple roles simultaneously. Does that complicate things further, with multiple preference sets in play?
Praveen Bodigutla (10:25)
There is difficulty, but at the same time, there is a nice carryover effect of preferences from one similar set of roles to another. The challenge is understanding how the data, preferences, and memory are organized. In the hiring assistant case, we have a natural tree-like structure for these preferences. A recruiter has multiple projects they’re hiring for, and each of those projects can have their own individual preferences at the leaf level, while broader recruiter-level preferences sit higher in the tree. This hierarchical structure lets us efficiently manage and retrieve preferences at the right level of specificity while still benefiting from shared signals across similar roles.