GLKB Investigate
When an agent thinks for 15 minutes.
Company
GLKB · Jie Liu Lab, University of Michigan
My Role
UX/UI Designer — sole designer on the product
Tools
Figma · Claude Code · React · Claude Design
Timeline
May – Aug 2026
Description
GLKB is a biomedical literature knowledge base. Investigate is its deep-research mode: ask one question, and an agent reads the literature for eight to fifteen minutes before it answers.
Context
I designed the eight to fifteen minutes a research agent spends thinking, from the only thing I had been given: the answer it produces at the end. A front-end developer built that design, and then the one after it, and both were wrong. The third time I recorded a live run and measured it myself. Investigate is live on glkb.org — the toggle opens without an account; watching it run needs one.

The product while it is thinking, not its answer. The four counts sit above the step list rather than beside it, so quantity and activity are read in one pass, and only one step is ever open — every finished stage has folded behind a single past-tense line instead of being cleared away.
I was the only designer on the product, and I designed Investigate end to end: the run interface, what it says at each stage, the report structure and the reference layer. The pipeline is the agent team's, the front end is Bing Han's, and the feature was the product manager's to shape.
I was handed one answer and the sentence "it's gonna run long"
That was the brief. A single finished report an internal run had produced, and a warning that the thing takes minutes rather than seconds. Everything about the product happens in those minutes, and neither of the two things I had been given contained any of them.
The audience makes the gap expensive. GLKB exists so that a language model's claims can be checked against citable literature, and biomedical researchers are the last audience on earth to extend credit to a screen that has gone quiet. Three things pulled against each other: the run takes minutes, and its real numbers arrive late and out of order; it has to live inside an existing chat column without becoming a second application; and on a product whose whole promise is that claims are checkable, the progress display cannot be the part that lies.
There was no in-house precedent, so I built one — the same question through four deep-research products, inside a wider benchmark of fifteen. None of the four solved the wait. One of them was useful theft: it publishes a retrieval funnel, but afterwards, as a figure in the finished report. Researchers already read a funnel as credibility. Nobody was running one live.

Redrawn, not captured — no competitor's interface is reproduced here. Each product is shown at the single moment it discloses the most, so the comparison is between disclosure strategies rather than between screens. The one borrowed idea is the retrieval funnel, and only one of the four prints one — after the run, as a figure inside the finished report. None prints it during.
Plot Twist 1: I designed nine minutes of waiting without ever having watched nine minutes
I read the answer, took it apart, and drew the wait as frames in Figma. The product manager asked for three additions — a notify-me, a progress bar, and a countdown — and I drew those too. Then it went to the front end and came back as a question I could not answer: it isn't just a front-end issue, the content and the steps are very different. The product manager looked at the build and said "yo that looks bad." I looked and agreed. The countdown never survived that round at all: I built it because it was asked for, and the agent team's answer was that it could not exist — they do not know how long a run takes, and there is no continuous quantity underneath a bar to fill.
That was the first thing I actually learned. This was not UI work, and it was not even UX work. I had been designing the shape of a process I had never seen. The countdown I drew is not what shipped. What shipped is an estimate, and nothing inside the run knows enough to make it a promise — it is the one number here that has to be allowed to be wrong.

Five frames sampled a minute apart across the first five minutes of the recording, differenced pixel by pixel. The frames themselves are not reproduced — what is drawn is the result, which is that no two of them differ at all.
Plot Twist 2: The second design was wrong for the opposite reason
So I asked the agent team for the chain, and they gave it to me. It was the first time I had seen what actually happens inside one of these runs — not just in this product, in any of them. Tool calls, the papers being read, every step. I took it apart and designed again, and the build came back and the product manager said "yo what happened." This time I could see why. The agent's own tool-call strings had reached the surface where a researcher would read them — in my own words, I under-specified what to show and what not. And underneath that was the real one: a long stretch of the run in which the interface had nothing to say at all.
The chain had solved the wrong half. It told me the order of everything and the timing of nothing — there are no timestamps anywhere in that trace. So I recorded a live run and annotated it, marking when each thing happened and what happened, and that recording is where the missing timing came from. With it I broke the run into keyframes, and asked the same three questions at every one: what has the front end been given, what is it waiting on, and what should it be showing. That method was mine. Not the product manager's, not the agent team's. The result is fifteen keyframes across ten stages. Two of its rows are not instructions to the interface at all — they are requirements on the backend, and they are in the pipeline now.

Two frames of one recorded run, seventeen seconds apart. The elapsed clock is the only thing that changes: 852 of 1,212,480 pixels differ between them, and none of them is an element of the interface. This is the stretch the second design produced.
The Outcome: what the product says while it is thinking
Everything below is live. The mapping is the specification the shipped interface was built from, and it is the only artifact here that is not a screen.

The answer surface during one real run: the question at the top, the four counts, the steps printed as each one finishes, and the reference panel held beside them rather than after them.
Keyframe Mapping: Specifying Time Without a Clock
Ten stages, T0 to T9, and fifteen keyframes inside them, each a row of six columns: the trigger that fires it, the stage mark, the four counts, the active step, and the step's content. I then built it twice, because a table cannot be argued with and a running thing can. One build plays the run on a timeline with its timing constants named, so a developer watches the pacing instead of inferring it. The other steps scene by scene, fourteen of them before the fifteenth reveals the answer, and prints the backend trigger beneath the state it produced, so the specification reads in either direction — from an event to what the screen must do, or from a screen state back to the event that has to exist for it.

The trigger is printed beneath the state it produced, at the same weight, so the specification can be read in either direction. The counters are drawn mid-flight — three landed, one still climbing — because a row in which every number has arrived is a row that cannot show the rule.
The Clarifying Question: Interrupting a Run Already Burning
The agent stops mid-run and asks you a question. I proposed it, and I won the product manager and the backend over to it; nobody had assigned it. On a materially underspecified query an ambiguity gate fires before the run commits and puts one to four scoping questions in front of you. The default is not to ask, the questions are skippable, and the finished report opens by recording the scope you chose, so an interruption leaves a trace in the output rather than vanishing into it. Every product I benchmarked either asks before you have committed anything, or never asks at all. Asking inside a run that is already burning is the interesting case: interrupting costs something, and so does spending eight minutes on the wrong reading of the question.
The shipped form loaded with the second option pre-selected in every group. A pre-selected answer to a question you are asking because you do not know the answer is not a default — it is the question answering itself. It was removed.

The question is placed inside the run rather than in front of it, so the interruption sits in the same column as the progress it is pausing. The options are the four facets this query was actually split into — the scoping decision the run is about to make on your behalf if you say nothing.
The Reference Layer: Making a Citation Cheap to Check
The reference system separates abstract-level evidence from full text — the distinction that decides how much weight a claim can carry. Each response filters to its own references, and they sort by citation count or by year. And any claim in the answer expands to the sentences it was drawn from: the inline marker opens a card carrying the source sentence itself, with the matching entry highlighted in the panel beside it. A reader gets from a sentence in the answer to the sentence in the paper without leaving the page, which is the only test of a citation that decides anything. One thing did not ship: an icon pair marking each reference abstract-level or full-text. A plain number marker shipped instead.

One claim expanded to the sentence it came from, with the matching reference lit in the panel beside it at the same instant. The decision is that the two are shown together rather than linked: a citation you have to navigate to is a citation most readers will not check.
The Content System: Saying Something True Before Anything Arrives
Some of what this interface shows is not measured. In my own words at the time: "a lot of the time, we actually don't know, we have to fake. Like the paper number ticking up, the progress bar." Every fake in the product is drawn against a single rule. The counter may under-report and it may never over-report. A number that has not landed still has to make it look as though work is happening, so it climbs on a curve that cannot reach the true value and locks to the real figure within three seconds of the backend reporting it. Until then the slot holds a dash rather than a zero — a zero is a quantity, and only the backend is allowed to state one. The same rule sorts the language: present tense while a stage is open, past tense the moment it closes, and nothing discarded — a finished stage folds behind its own summary line instead of scrolling away.
The estimate is the one number here I would take back. Every other number is permitted to fall short of the truth and catch up; a time that counts down cannot do that, and it is drawn against the same rule anyway.

Four counts drawn at the four separate moments they arrive, never together. The empty slot holds a dash rather than a zero, because a zero is a quantity and the backend has not stated one — and the curve beside them is drawn so that it cannot reach the true value. Under-reporting is a drawing decision before it is a rule.
Impact
Investigate is live on glkb.org, where the platform is in use by more than fifty researchers across Michigan and four outside institutions. One real run: 1,341 papers retrieved → 53 screened → 20 extracted → 17 cited. The progress interface, the counters, the mid-run clarifying question, the report structure and the reference layer are all live. The notification work is mostly in: the in-run control and the completion email both shipped, and the settings row did not — and the difference is that what shipped lived inside the flow I had specified trigger by trigger, and the settings row was a frame.

The same status block at the four stages of one run. Each count lands only when its backend step reports it: 1341 retrieved first, screened climbing to 53, then 20 extracted, and 17 cited last. Until a number arrives its slot holds a dash, not a zero.
Nobody argues with nine minutes of silence they have all just sat through
I had asked for time with the agent team before any of this and been told everyone was too busy — work on it yourself first. So the recording is what ended the argument, and it did not end it by being clever. We got on a call and went through a live run together, in real time, all the way through the stretch where nothing happens. The method is what got the meeting.
![A screenshot of the GLKB Investigate completion email as it was received: from GLKB noreply, subject [GLKB] Your investigation is ready, carrying the heading Your investigation is ready, the question cut to two lines, a View full report button, the completion time, and a result card printing 6921 retrieved, 52 screened, 26 extracted and 21 cited. Beside it the sender line and the result card are enlarged, over a list naming what arrived and what moved from the design.](https://framerusercontent.com/images/tfQ7hPb4Sz8GzZpHMUk5UFMlQ2I.png?width=1840&height=1380)
The run-completion email as it arrived, with the sender line and the result card enlarged beside it. A reader who leaves gets the four counts back in the order they landed, with one button to the report they were waiting for.
I specified what to show and never specified what to hide
Two rounds of this went to implementation before any of it was right, and Bing Han built both of them. The first came off drawings I had made from a finished answer — I had never watched the thing run, so what I handed him described a screen and not the minutes it had to survive. The second came off the reasoning chain, which was better and still wrong: I under-specified what to show and what not, and the agent's raw tool-call strings came through onto the surface where a researcher would read them. I did not draw those. I left the door open for them, and he built exactly what I gave him, twice.
Credit
Product manager — Zhiyuan Liu
Agent team — Xiang Zhang
GLKB — Zhaowei Han, the PhD student who led GLKB
Front end — Bing Han
Lab — Jie Liu Lab, University of Michigan