Will Held is a member of technical staff at Open Athena’s San Francisco office, where he works on the Marin project, a general-purpose large language model developed entirely in the open. In this Q&A, he looks at how research in computational linguistics evolved into today’s large language models, the tradeoffs of working in the open, and the English language’s linguistic quirks.
What’s one of your earliest experiences with AI technology?
Will Held: When I showed up at university for undergrad, I was talking to a professor in computer science—Professor Nizar Habash—who did computational linguistics for Arabic, and he had just started his lab at the university. I had studied Arabic all through high school, so I got really lucky that I could come in and learn from him directly, before he had a ton of research assistants hired. That was probably pretty far from how we define AI today—in those days, it was statistical language processing, and we were trying to bake linguistics into the algorithms. But even at that time, it was starting to be the very early forms of what are now language models.
And for a course that Professor Habash taught, I did a project on learning word embeddings for Arabic. People were already doing that type of stuff, but I got exposure to taking a big corpus of text and, by learning how to predict co-occurrences between words, doing some interesting things about learning what the meaning of a word is. While that is very, very different in terms of scale and engineering, it is the underlying statistical process that we still are still using for all of the stuff that we do.
So that was my first foray into the field, and very happenstance and lucky. I thought, “I like computers, and I have been studying Arabic, so I guess I'll do some of this computers and Arabic stuff,” and it has spiraled from there.
What drew you to work at Open Athena?
During my PhD, I had worked on a couple of the Llama models during internships, again through happenstance. I'd been hired to do machine translation at Meta, and then they were like, “We're gonna start putting more resources into training language models.” So I learned a lot at Meta under the really early open-source Llama stuff. And I really enjoyed how much I learned through that experience of developing open models, but also I really enjoyed seeing the output of that, and how much research was created on top of the Llama models.
And so when I came back for my last year of my PhD, Percy Liang and David Hall reached out to me and said, “We're doing some open source stuff in academia and have a lot of computing resources compared to a lot of academic efforts.” I started working on the Marin project through them. And then it became clear that I could continue doing that as my full-time job after my PhD, through Open Athena. Even in a time of a lot of unique research opportunities, it still felt unique to be so well-resourced, even at a non-profit, to try to continue advancing open science.
In general, the parts of academia that I love are trying to teach people, and talking about your work, and trying to make your work digestible too, both to the broader field but also to audiences outside of just your field. And so Open Athena was all of the elements of academia that I like, doing research and trying to disseminate it. But it’s in a new format that addresses some of the elements of academia that are very different if you become a faculty member, where there's a lot more logistics and more of a people-management job, rather than a research job in many ways. So it was a really good fit for me in that I still got to focus on driving good research, but unlike at a closed lab or something like that, I still got all the elements of education and dissemination of knowledge being a really core part of what I was working on.
What sort of non-project-specific work do you do at Open Athena?
For better or for worse—I don't know if the science foundation model colleagues like this or dislike this about me—when I joined, I was really excited about trying to find ways that we could transfer more of the language model science and the best practices that have been learned through, you know, an absurd amount of research expenditure on language models, to be transferred to and useful in other scientific fields.
One of my formative experiences as a researcher is, I started out more interested in the language part. I was in a computational linguistics lab, and occasionally this makes me feel very old, but I've been doing NLP-adjacent things for about a decade. And the trajectory over that decade has been, we used to care a lot about really understanding the language, and understanding the mechanisms, and knowing about the linguistics literature for a particular language. While I think there's still a lot of value to that and enjoyment in that, the field has actually made progress largely through moving away from that, and letting the bitter lesson of “more computation, less assumptions” work its way out.
The bitter lesson is this idea that started in reinforcement learning—chess, Go, and general games—which was, “approaches that just make better use of more computation, with less constraints, tend to do better than approaches that make more assumptions and might look very good at small scales, but don't actually scale as well with more compute.” For many years in chess AI, they did all this, “oh, we're gonna encode how experts think about chess.” And ultimately what made superhuman chess engines was just basically doing search at an incredibly large scale and not encoding anything about how human experts consider it. Now the bitter lesson is very much a part of language model discourse, but it started in pre-modern AI.
And certainly for language, that's been the trajectory too. We make fewer and fewer of these linguistically-informed assumptions, and make more and more use of compute. So when I joined Open Athena, I was probably an annoying agitator saying, “Well, we're looking at these different models, and there's a lot of different assumptions they make from different scientific disciplines. Is it worth—at least as a straw man—trying to think about making fewer assumptions? And in making fewer assumptions, can we leverage more of the stuff that we've learned at really large scales from language modeling science?”
I don't know if this has yielded any benefits yet, but I spend a lot of time chatting with my scientific colleagues about, this is how we do this in language modeling world, is this transferable to your domain? And I hope there's value to the science—certainly there's a lot of value to seeing if something does transfer to a totally different distribution of data, to DNA or to proteins. Then that tells us that it's something fundamental about the learning processes of these models rather than the data distribution, whereas if something doesn't transfer, then we're like, “well, this is something that's somewhat language-specific.” And so I spend probably more time than I should, if I was maximizing productivity, trying to figure out how much we can transfer over and how much we can transfer back, across Marin and some of the other scientific projects.
What do you like best about working here?
There's what academia is in the ideal, and then there's the reality of what academia is. The reality of academia is that a lot of making science happen is getting together the funds to fund the right experiments, and doing that involves building brands and sort of making yourself famous. Academia is a field where you are the brand to some extent, and so there's a lot of brand-making in yourself.
And one thing I really enjoy about Open Athena is—as ironic as it is, in this context of an interview where we're gonna put out a blog post—I think we're a culture that is, very much so, more interested in trying to figure out how we can contribute to a particular space than in necessarily building our own brand up. And it feels like everyone is pretty aligned on pursuing useful knowledge and disseminating that, even if doing that in certain ways doesn't necessarily accrue us the brand recognition.
In all of the different open development stuff, I think doing open development is a very irrational thing from the perspective of brand building. You're constantly disseminating knowledge, and you don't get this big moment where people feel you've released a huge update that changes everything. Instead you're just constantly diffusing the knowledge. And I think that speeds everything up, and hopefully means that people can learn quicker, but it is something that only makes sense if you're really trying to maximize the rate of diffusion. It means that you might not be the person who gets credit for some big step forward, because you're constantly sharing everything you know, rather than saving it so that you can share it all at once and make it appear that there's a big leap.
So that element—that we really care about driving progress, even if that means doing it in ways that probably are strategically difficult for our own brand building—is something that I actually really enjoy.
What does a typical work day look like for you?
As experiments get more expensive for language modeling—or just take more time even if they're not expensive in dollars—I think a lot about trying to sequence out my days in terms of timescales. There’s four timescales I think about a lot, which is, what can I get done in an hour? Eight hours? A week? And then, what can I get done in a month?
So for me, what that means is, the eight hours matters because it's something that I can run before I go to sleep, and I can come in the morning and get some big updates on. Usually the first thing in the morning is coming and checking in on what’s finished overnight—hopefully, all of my jobs have run successfully, and all of my experiments have come back without errors. Then I figure out how that updates different bits of what I'm working on.
I try to design my day around what I learned from the previous night’s experiments. What are the things I need to get done in hour chunks, so that by the end of the day, I'm ready to launch another set of things that's gonna take that eight hours, that I can come back to the next day? If I end a day and I don't have a set of eight-hour experiments that I'm gonna come back and get information from, it feels like I've wasted an evening, because the computers weren't working hard.
And then I’m also structuring my day around what little meetings and research discussions I have with colleagues, to figure out what exactly I need to get done to launch that next eight-hour thing. And then, similarly, when I come in at the beginning of the week, ideally I have something that's been running the whole week, and so Mondays, I'll spend more time just looking at experiments that have finished over the course of the week and the weekend.
Day by day, things can look quite different in how I'm spending my time, because to me, so much of doing science in a good way means that you should be updating your plans very aggressively based off what new signal you're getting. And so I think more about those little timescales of that hour, eight hours, week, month. That means, day by day, I'm trying to update what I'm gonna do to maximize for each of those timescales.
Do you have any favorite AI tricks or tips, whether for home or for work?
Both for work and for home, in this era of agents being able to do stuff, I bought myself a little Raspberry Pi that I can just have sitting in my house. I use that a bunch for work, in that it's where all of my stuff lives and runs, launching jobs on different clusters, but it also runs my coffee machine. It turns my coffee machine on and tracks the power usage.
Ironically, in the age of AI and agents, having a desktop or something that's always on is more important than a laptop. I think too many people live their lives afraid to close their laptops, because people are so worried about their Claude Code or whatever session dying. I'm like, “No, just buy a computer and set it up somewhere so it's always on.” That way you can live your life without concern about whether your laptop's open. And it means that, for me, the actual computer that I have is not very meaningful, because everything's actually running somewhere else. So if I forget my computer, I can use my phone, and I can connect to the machine back at home.
That's something that, to some extent, I've always done, but I think it is even more important in the age of AI. In many ways, I think AI has reinforced things that I've learned from other experimental science, which is, try to make sure a computer somewhere is always working. I think that's most of what I learned in my PhD: you don't necessarily have to work super hard, but the computers always need to be working very hard. But it takes a lot of thinking to make sure that you can actually structure that.
I guess my simple tip is, buy a Raspberry Pi or a desktop. In the age of AI, laptops are dead.
Do you have a favorite piece of science trivia?
Even before it became the global language, English as a language was already essentially a pidgin between Latin and Germanic languages, and this reflects itself even at very low levels. So the French essentially conquered the Britons, and when they did that, French became the higher-class language, and English’s Germanic roots all became lower class. And how we reflect that in the modern day is that, for most languages, the words you use to refer to an animal, whether you're eating it or whether it's alive, are actually pretty close. But in English, we use the French root when we're referring to an animal that we're going to consume. So, beef is boeuf, we have the French root, Latin root, for the consumed form. And then we use the Germanic form for the animal when it's alive and messy in the field: “cow” comes from the Germanic root. And so you get this reflection of the fact that the French was upper class, so their animal words were used for the refined setting of consumption.
And this is probably folk linguistics, but it’s claimed to be the origin, in part, of the phrase "Pardon my French." The British were saying, “oh, if I use a crass word, you can ironically say ‘pardon my French,’” because in fact, French was the upper-class phrasing.
Can you tell me about a hobby or activity you enjoy outside of work hours?
I'm a big runner, less so now, but throughout my PhD especially. I've run three marathons and one ultramarathon. It was 50 kilometers, outside of Huntsville, Alabama, basically through the frozen woods.
I like getting outside and covering lots of distance just to see things, maximize how much I'm seeing in a day, and get a better sense of place. Now, in the Bay Area, I've been getting more into biking, because there's better bike infrastructure and it's safe. I find if I'm in a car, I don't actually know where things are very well, whereas running or biking gives me a much better sense of place. So I did most of my running when I was in Atlanta, Georgia for my PhD, and there are sections of Atlanta that I know incredibly well, just because I've covered it all by foot, whereas anywhere that I've driven more, I don't retain.
Do you have a signature response emoji, or one you use most often?
Probably at Open Athena, I would guess the salute emoji. This is something I pulled more from my old lab culture: I don't know how it was formed, but anytime our PI would send a message about some new thing we were doing, or something like that, we would all respond with the salute emoji. And so I still use the salute emoji a lot, which I think is probably the emoji I use that’s the most out-of-distribution from general use.
Cite this post
@misc{bushwick2026_meet_our_team_will_held,
author = {Bushwick, Sophie},
title = {Meet our team: Will Held},
year = {2026},
month = {aug},
howpublished = {\url{https://www.openathena.ai/blog/meet-our-team-will-held/}},
note = {Open Athena Blog}
}