Research
Last updated October 3, 2026.
Wolf is built on published research. This page lists the work behind each part of it, with a plain note on how it shapes what Wolf does. Where something is still planned rather than built, we say so.
Memory and understanding
- Generative Agents: Interactive Simulacra of Human Behavior. Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang and Michael S. Bernstein. 2023. Agents that keep a record of experience, synthesize it into higher-level reflections, and use those to plan.
How Wolf uses it: your Wolf keeps a short, current picture of what's going on in your life (goals, open threads, the people involved, what's coming up) and rewrites it as things change, instead of piling up raw notes. - MemGPT: Towards LLMs as Operating Systems. Charles Packer, Vivian Fang, Shishir G. Patil, Kevin Lin, Sarah Wooders and Joseph E. Gonzalez. 2023. Treats a model's context like computer memory, with tiers that are always loaded and tiers brought in when needed.
How Wolf uses it: the rules Wolf follows on every message stay small and always present, while the detailed steps for longer tasks (finding a time with a group, splitting a bill) are opened only when a request needs them.
Acting with tools
- ReAct: Synergizing Reasoning and Acting in Language Models. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan and Yuan Cao. 2022 (ICLR 2023). Models that alternate reasoning with actions that fetch real information, which reduces made-up answers.
How Wolf uses it: for live facts (hours, scores, weather, your calendar) Wolf looks them up with a tool first and answers from what it found, rather than from memory. - Lost in the Middle: How Language Models Use Long Contexts. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang. 2023 (TACL). Models use information at the start and end of long inputs better than the middle, and do worse as inputs grow.
How Wolf uses it: we keep Wolf's standing instructions short and put the most important rules first, and we measure reply quality and speed whenever they change.
When to speak up
- Principles of Mixed-Initiative User Interfaces. Eric Horvitz. Proceedings of CHI 1999, pages 159 to 166. An assistant should act on its own only when that's worth more to you than doing nothing, and should weigh the cost of interrupting you against waiting.
How Wolf uses it: when Wolf reaches out on its own, it has a high bar, a daily limit and quiet hours, and anything you mute stays muted. Reminders you asked for arrive at the time you chose. - Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance. Yaxi Lu and colleagues. ICLR 2025. Measures proactive help against whether people actually accept it.
How Wolf uses it: proactive messages are judged by whether you'd want them, so Wolf stays quiet unless one thing is clearly worth your attention.
Privacy between people
- Privacy as Contextual Integrity. Helen Nissenbaum. Washington Law Review 79, page 119. 2004. Privacy means information flows that fit the norms of their context.
How Wolf uses it: finding a time with friends shares only whether you're free or busy, never what your plans are; group results show counts ("9 of 10 can make it"), never who can't; family chat answers come from a separate answerer that can't see anyone's private information. - Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory. Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri and Yejin Choi. ICLR 2024. Shows models reveal private information in contexts where people wouldn't.
How Wolf uses it: we don't rely on the AI's judgment for what crosses between people. Fixed programs, not the AI, compute free times, build the messages one Wolf sends another, and decide what's shared. - PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action. Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu and Diyi Yang. NeurIPS 2024 Datasets and Benchmarks. Finds that agents leak sensitive information when acting, even when told to be careful.
How Wolf uses it: messages between Wolves use fixed formats with plain-word fields only, every send to someone new needs your yes to the exact message, and replies like "yes" to a connection request are handled by the server, not the AI.
Security against hidden instructions
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz and Mario Fritz. 2023. Shows that instructions hidden in emails or web pages can take over an AI assistant.
How Wolf uses it: everything Wolf reads in your email, calendar, the web or a family chat is treated as information, never as instructions, and Wolf can only run the specific commands on its own list. - AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer and Florian Tramèr. NeurIPS 2024 Datasets and Benchmarks. A standard way to test agents against these attacks.
How Wolf uses it: we attack Wolf ourselves with planted fake secrets and hidden instructions, and fix what we find before adding features. - Defeating Prompt Injections by Design. Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis and Florian Tramèr. 2025. Separates what you asked for from untrusted data, so that data can't change what the system does.
How Wolf uses it: actions that involve other people run through fixed steps that untrusted text can't redirect. Planned next: a check-only mode for Wolf's background reviews that can't send or change anything at all.
Suggest a paper
Know research we should read? Email support@moriahlabs.com.
