ScienceJury

Multi-agent feedback system built to consolidate AI-driven workflows for academic writers.


Timeline

October – December 2025

Team

1 Project Lead

1 Principal Investigator

2 UX Researchers

1 Software Engineer

1 UI/UX Designer (me!)

Venue

CHI 2027 (Submission pending...)

Project description

ScienceJury replaces the context-switching and window-splitting experience of using AI for writing feedback, providing a structured, specialized document-integrated review experience.

What I owned

• Mapped out the interaction model, information architecture, and user journey.

• Assisted researchers in translating research findings into design goals.

• Design prototypes for our experimental and control systems.

Impact

+84% more unique critique points vs. single-agent baselines.

+18% alignment of AI-generated feedback with real human peer reviews.

+24% increase in draft improvements reported by real authors.

Problem

Using existing AI tools for feedback forces writers to choose between context and range of perspective.

Editor-embedded AI anchors feedback to your text, but gives only one perspective.

Prompting multiple perspectives gets you range, but nothing stays connected to your draft.

Core design challenge:

How do we build an agentic feedback system that offers multiple perspectives and stays anchored to the text it references, without it collapsing under the weight of its own output?

Formative study

Does prompting multiple perspectives actually result in better critique?

Our team compared the critiques of single LLM agent against a curated, three-agent panel on 50 real papers with real human reviews. We measured each framework’s similarity to real human reviews and uniqueness in the issues surfaced.


The multi-agent approach identified nearly twice the unique issues that are more closely reflective of what human experts flagged.

Multi-agent
Single-agent

Human alignment (by cosine similarity)

Multi-agent: 0.45
Single-agent: 0.38

Unique issues covered (average % per paper)

Multi-agent: 33.4%
Single-agent: 18.2%

But without structure, more perspectives create more noise.

Increased volume and diversity of feedback makes it difficult for authors to prioritize and act on suggestions.

Not all feedback is equally actionable or relevant to the author’s intent.

Conflicting critiques require authors to interpret and reconcile feedback from multiple sources.

This formative study gave me a key insight for the system’s design:

More feedback volume requires heavy design investment in how the feedback is organized and surfaced.

Design strategy

I designed ScienceJury to support authors in navigating, interpreting, and applying feedback from multiple perspectives.

Support multi-perspective critiques through customizable and specialized reviewers.

Let authors debate, question, and reason through feedback instead of passively receiving it.

Reduce context-switching by anchoring every critique to the exact sentence it refers to.

Solution

Configuring the agent panel (01)

ScienceJury suggests a starting panel of three reviewer profiles, which authors can customize, regenerate, or remove.

Every agent shows its rationale, review focus, and cited scholars, so authors can judge the feedback instead of just accepting it.

Agent configuration flow


Interacting with agent feedback (02)

Feedback is broken into sentence-level items, anchored to the exact passage, and color-coded by agent, so authors can see where perspectives converge or conflict.

Each feedback item opens into a thread, ask follow-ups, push back, or adopt a suggestion inline, without leaving your draft.

Feedback interaction flow

Results

In a between-subjects study with 22 academic authors, ScienceJury positively influenced how they edited their drafts.

Participants revised their in-progress papers using either ScienceJury or a single-agent control. The control condition had the same editor and underlying model, but only one agent in a more familiar LLM interface.


Feedback rated significantly more specific, diverse, and aligned with expert review.

Authors reported deeper engagement with their own drafts.

+84%

more unique critique points vs. single-agent

+40%

higher specificity rating vs. single-agent

4.7 / 7

average draft improvement vs. 3.8 in single-agent

Every suggestion and review or comment I leave on my document has potential for a whole conversation to spring from it and has context that I can’t capture with just one comment.

But [ScienceJury] uses the AI to capture all that context per sentence, which is really helpful.

– Participant 8

[Deconstructed feedback] gave me multiple doors to enter the argument again. Each agent pushed me in a different direction, so I could decide which made the most sense for the story I want to tell.

– Participant 7

However, ScienceJury didn’t fully solve for over-trust.

Reviewer personas written to sound credible and domain-specific gave some participants an outsized sense of authority.

It felt authoritative when it framed things like ‘As an HCI expert, I believe. . . ’ It made me consider the feedback more seriously. But I am also worried I focused too much on this review.

– Participant 10

This feedback tracks with existing research on how the source framing of the feedback shapes engagement (human vs. AI-generated).

Credibility and trust should have been framed as design problems, rather than side effects of good persona-writing.

This opens up a question for future iterations:

How do you make AI feedback feel credible enough to be useful, without it becoming so authoritative that it overrides the author’s own agency?

If given more time, I’d explore softer persona framing, explicit uncertainty signals, or confidence indicators surfaced alongside each feedback item.

Constraints and pivots

What forced us to adapt, and what I learned from it.

This project ran from October to December, leaving us with two months to design, build, and run preliminary user studies. Our project lead needed the paper to defend her PhD thesis before starting an industry job. That timeline shaped every decision about the system.




Reflection

Insights and lessons I’m taking away from designing ScienceJury.

Under a tight timeline, feature scope should be cut to what's testable in the sprint.

We cut named personas and extra features to protect what the system needed to actually function and test within the timeline. If this were to be monetized or pitched as a start-up concept, a longer list of features would be necessary to demo and sell the idea better.

Where feedback lives changes what people do with it, not just how fast they find it.

Writers stopped padding around weak passages and started replacing them directly once the critique sat next to the sentence it was about. The interaction logs surfaced that shift, even though I hadn't directly designed for it.

Keep exploring

About
Archive
More case studies coming soon