The evidence panel I argued for did not ship as designed.
At Margin Research, I designed a way to inspect generated findings. The launch decision exposed a disagreement about what a trustworthy first impression should look like.
Folio at Margin Research.
Folio was Margin Research’s new product for searching interview repositories. I joined as a fractional design lead while the team already had a working retrieval service. My brief was to turn an impressive answer demo into a review workflow a researcher could stand behind.
What I owned
I defined the review workflow, built the interaction prototype, co-authored the evaluation rubric with Inez, and reviewed implementation. I did not train the model or own retrieval quality.
What I worked with
A working answer-generation service and a source library. The initial direction treated citations as a finishing detail on an otherwise complete answer.
The people I worked with
- Daniel Cho
- Founder and product lead · launch decision
- Inez Duarte
- Retrieval engineer · source quality and evaluation
- Leah Frost
- Research operations · study corpus and recruitment
- Owen Bell
- Frontend lead · citations and keyboard behavior
A fixed launch date and a leadership preference for a cleaner first impression. I could influence the release, but did not have the final call on the default panel state.
A citation looked more certain than the claim.
The first prototype could produce a fluent answer in seconds. It could not reliably help a researcher distinguish a supported insight from a plausible interpretation. That distinction was the product’s central trust problem.
The archive mixed current research, old studies, small samples, and notes without clear provenance. A confident summary could flatten all of them into a single voice.
The brief was to improve the chat experience. I reframed it around the review workflow: how does someone move from an interesting answer to a defensible claim they can share?
Can a researcher tell when a plausible answer deserves to be rejected?
Give a bad answer somewhere to show its work.
Inez and I reviewed sixty answers against their source passages. We scored support at the claim level: direct, partial, contradictory, or missing. A syntactically valid citation was not automatically useful evidence. We disagreed on nine of the first twenty answers, then rewrote the rubric before evaluating the remainder.
I tried a full answer with footnotes, source cards underneath, and an adjacent passage panel. The panel made it possible to inspect a claim without losing the question. It also consumed the space the founder wanted for a clean, fast answer.
The prototype treated a citation click as selecting a claim and focusing its passage. Keyboard users could move between claim and source, then return to the same place. Conflicting passages remained separate. “Not enough evidence” offered a change of study scope instead of another confident paragraph.
Review a claim at the level of its source.
Claim-level source focus
Selecting INT-07 opens the exact excerpt supporting that claim. A second source replaces the focused passage without losing the original question.
Disagreement stays visible
Solo participants and workspace administrators can describe different barriers. The interface preserves both instead of averaging them into consensus.
An empty answer is a recoverable state
The product explains that the selected studies contain insufficient evidence. Searching all studies is an explicit scope change, not a silent expansion.
Why do teams hesitate to invite colleagues?
Invitations ask for commitment before value is clear.
Larger teams need access control earlier.
“Inviting somebody feels like I have committed to using it.”
We deliberately put weak claims in the test.
Eight researchers completed counterbalanced evidence-finding tasks in the existing search and the review prototype. We inserted weak citations deliberately. After launch, a separate usability review with six researchers examined the collapsed-panel default; the groups and tasks were not directly comparable.
Four of eight researchers accepted a weak cited claim in the original flow.
Revision. Made the supporting passage visible beside the selected claim.
Readout. 7 of 8 caught a weak citation in the prototype.
The prototype kept study scope clear but made answers feel visually dense.
Revision. Shortened excerpt previews while preserving claim and source focus.
Readout. Median evidence-finding time moved from 13 to 7 minutes in the prototype tasks.
The launch version collapsed the panel before the first citation interaction.
Revision. Retained source markers, a persistent source count, and one-action passage access.
Readout. Only 3 of 6 opened a source without a prompt in the separate post-launch review.
The release was less protective than the prototype.
Two different studies. These are not a before-and-after comparison.
caught a weak citation
Prototype; seeded usability task
opened a source unprompted
Separate review of the shipped default
Daniel chose to ship the evidence panel collapsed by default. The rationale was a cleaner first screen and a faster perceived response, particularly in sales demonstrations. I disagreed because the prototype’s strongest finding depended on seeing the passage without first deciding to inspect it.
We agreed on a smaller set of protections: stable claim markers, an explicit study scope, preserved source focus, and an honest insufficient-evidence state. I recorded the lost default visibility in the launch decision and proposed a follow-up experiment with both panel states.
In the six-person release review, only three researchers opened a source without being prompted. That was a concern, not a causal comparison with the earlier eight-person study. The A/B test had not run when my engagement ended. I would not attach the prototype’s 46% task-time improvement to the shipped product.
I would agree with leadership on the review behavior we needed to protect before polishing the first screen. We shared a word, “trust,” while measuring different things: the confidence of a demo audience and the skepticism of a working researcher.