AI as a Search EngineA year ago, I would have told you that AI makes for a pretty crappy search engine.
In fact, I did tell you that! In this newsletter!LLMs are notorious for hallucinating information. So, if you’re not careful, you might find yourself scheduling a pre-interview with a woman who ChatGPT says negotiated a major humanitarian deal in Sudan…only to discover that she never even worked in that country.
Yeah, that was embarrassing.But sometimes you have a question that’s just too complex to drop into Google. And over the past year, LLMs have gotten much better at finding sources to help you answer those questions.
Notice what I said there? I didn’t say that LLMs were great at actually
answering my questions. I don’t use AI to give me answers.
I use AI to help me find reliable sources. Just like I use a search engine.
An example. Just last week, I was searching for a celebrity-type person who has experienced a very recently identified mental health phenomenon.
But since this phenomenon was recently named, I couldn’t just search for the name of the phenomenon itself. I had to search for the symptoms of the phenomenon.
This is not a great task for a traditional search engine. You’re going to end up with too many irrelevant results.
But to my great surprise, in just a few minutes, ChatGPT came back with a link to an article about a well-known actor who experienced this phenomenon right before taking the stage during the promotional tour for his latest blockbuster movie. The symptoms he described exactly matched the symptoms of this phenomenon, even though he didn’t use the same words to describe what he had experienced.
ChatGPT also returned some suggestions that weren’t a great match. Which is why I use AI as a search engine, not an “answer” engine!
The prompt I used:Find celebrities or well-known people who have spoken or written about suffering from [name of phenomenon] or who have talked about experiencing a sudden onset of symptoms like X, Y, and Z after experiencing [inciting event]. Cite your sources.Cite your sources is key. Because I’m not looking for answers. I’m looking for sources.
Why I don’t stress about using AI for this task:These days, even if you’re just using Google, you’re using AI. So there’s nothing inherently less ethical or more environmentally damaging about using Claude or ChatGPT. If anything, you’re getting to the correct result more quickly, which limits the amount of searches that you do.
AI as a Dumber DownerMy producing partner and I recently wrapped our second longform investigative project. And wowza, this one was a doozy.
We’ve collected dozens of hours of interviews, hundreds of newspaper articles, thousands of pages of court documents and 31GB of videotaped court testimony.
And there’s this thing that happens when you’re trying to make sense out of an overwhelming amount of information. Your brain gets so full of all the little details that you start to have trouble seeing how it all fits together.
Is that just me? Please tell me that’s not just me!In this case, my nemesis was the episode outline. I rewrote the damn thing so many times, I thought I was going to lose my mind.
Each iteration of the outline got better, but it also got longer. And after a while, we could no longer see the forest for the trees.
So. Many. Trees.Finally, our editors asked us to pare it way, way down and put it into a specific format.
But by this point, the idea of writing another version (my sixth!!) of this damn outline just made me want to cry. So I decided to upload the existing outline to AI, add in the feedback from our editors, and see what happened.
And … it worked!
The AI generated outline was far from perfect. AI didn’t always know what details to keep and what to leave out. And sometimes it made assumptions that were just plain wrong. So I ended up rewriting quite a bit of it.
But AI took the task from something that felt impossible and turned it into something manageable.
The prompt I used:I literally just copy-pasted the request we got from our editors and asked AI to do it. I’m not going to copy-paste that feedback in here, because that’s between us and our editors.
I did have to experiment a bit on this. I hated the results I got from ChatGPT, so I used a different tool instead.
Why I don’t stress about using AI for this task:I’d love to tell you that I had a well-reasoned, ethical argument for why I felt like this was a proper use of AI. But in that exact moment, I just needed to get it done.
AI as a SynthesizerSo, this particular longform investigation centered around a four month trial that took place more than 25 years ago.
We searched everywhere for the trial transcript, but official copies were destroyed more than a decade ago. And it took almost two years for us to get access to the actual footage of the trial.
By that time, we were deep into episode production. We knew the story we planned to tell, based on interviews and contemporary newspaper reports. But there was a lot that was contradictory. Or missing. Or just plain wrong.
So in the end, I had less than a week to review four months of court testimony to…
1. Confirm (or disprove) details provided by sources.
2. Fill in the blanks with the details no one remembers.
3. Find compelling audio to support the series.
Clearly, I could not watch four months of testimony in a week. It was simply not possible. So here’s the somewhat convoluted process I used instead.
(Note: If you have much shorter files, you can do this all in one step – just by uploading everything to the paid version of Gemini Notebook. But since some of my files were more than four hours long, that just wasn’t an option here.)
Step 1: Upload the video to Descript and generate a transcriptI used Descript because my client has an unlimited Descript account. And also because we were building our assemblies in Descript, so our audio was going to end up there eventually anyway. But if I were paying by the hour for transcription, I would have pursued a less expensive option.
Step 2: Upload the Descript transcripts to ChatGPT and generate speaker labels.If I had tried to go through and identify each speaker manually, it would have taken me all week. So instead, I asked ChatGPT to use context clues to identify each speaker and give me a time code for the first time they appear in the transcript.
This wasn't perfect, but the time codes made it easy to go back into Descript and assign names to speaker labels. (After first confirming that they were correct!)
Step 3: Upload the completed transcript to Gemini Notebook and start querying!The cool thing about Gemini Notebook is that it’s only looking where you tell it to look. So when I asked it about the trial, it was literally only using the transcripts I had uploaded as source material.
The prompt I used:I just asked questions using real language.
Here’s a good example. The defense lawyer told this great story about a feeding pump that malfunctioned during a demonstration by a prosecution witness. The prosecutor had no memory of this happening.
So I asked, “Did a feeding pump malfunction during testimony about the reliability of feeding pumps?”
Within seconds, not only did I have an answer to my question (yes, it did) but it also linked to the transcript, showing me the exact time code of the indicent. I then hopped over to Descript, where I was able to watch the whole incident from start to finish, to make sure that we were putting it in the right context.
Why I don’t stress about using AI for this task:I wasn’t asking Notebook to analyze the trial itself. I was using it to help me confirm facts that might otherwise be impossible to confirm.
I left the trial analysis to the humans.
AI as a Research LibrarianIf you’ve ever worked on a longform investigation, you know how it goes. You start on Day 1 with big organizational plans. Nicely nested research folders. A solemn vow to catalog every document, news article, or archive audio/video element.
But somewhere between Day 2 and Day 600, it all goes to shit.
By the time we reached the fact check phase, my brain had reached overload. I could remember what I had discovered, but had a much harder time remembering where I had discovered it.
Once again, Notebook came to the rescue. It made quick work out of hunting through tons of data and finding what I was looking for.
The prompt I used:Again, I uploaded the specific documents I wanted Notebook to review, and real language queries were the key.
In the end, I wasn't just confirming that my memory was correct. I was matching court documents against our sources' memories, to make sure our reporting is as accurate as possible.
Why I don’t stress about using AI for this task:Look, there were no privacy concerns here. I was literally searching through publicly available court records.
And if CTRL-F could have answered this question, I would have just gone with that. But with thousands of pages to search through, it was impossible to find a single search term that would return a reasonable number of hits.
Real language queries were my only hope.
So, where do I draw the line? Some
news organizations and
podcast production houses are developing AI policies. And I really applaud the effort.
I have my own rules. And like most of the things I do, I’m keeping them simple.
Karen’s Rules for Working with AI- Never trust AI. Ever.
- Find the right AI tool for the job. (I highly recommend Jeremy Caplan’s Wonder Tools, if you want to expand your options.)
- Use AI as a tool, not a service.
A little explanation on that last one. It comes from a conversation I had recently for
Building 32, a podcast I’m hosting for MIT CSAIL Alliances. (CSAIL stands for the Computer Science and Artificial Intelligence Laboratory – so these folks know what they’re talking about!)
The researcher I was interviewing is an architect who’s developing more sustainable ways to build things. And she believes in using AI as a tool. For example, as a way to quickly analyze huge data sets.
But she doesn’t want to turn over her agency and creativity to AI. So she makes sure that all the actual decisions are still hers.
It’s a tool. Not a service.Seems like a pretty good guideline to me!