Discovering capabilities under real world constraints.

Reading Between the Lines: The One-Sided Conversation Problem

Victoria Ebert1 Rishabh Singh1 Tuochao Chen1,3 Noah A. Smith1,2 Shyamnath Gollakota1,3
1University of Washington    2Allen Institute for Artificial Intelligence    3Hearvana AI
One-Sided Conversation teaser figure
We introduce the one-sided conversation (1SC) problem: making inferences from only one side of a conversation transcript. We focus on reconstruction of the missing content and creating summaries of the whole one-sided conversation

Abstract

Conversational AI is constrained in many real-world settings where only one side of a dialogue can be recorded. We formalize the one-sided conversation problem (1SC): inferring and learning from only one side of a conversation. We study two tasks: (1) reconstructing the missing speaker's turns and (2) generating summaries from one-sided transcripts. Evaluating models on MultiWOZ, DailyDialog, SpokenWOZ and Candor with both human A/B testing and LLM-as-a-judge metrics, we find that additional context improves reconstruction, and while large models generate promising reconstructions with prompting, smaller models require finetuning. Further, high-quality summaries can be generated without reconstructing missing turns. We present 1SC as a novel challenge and report promising results that mark a step toward privacy-aware conversational AI.

Demo Videos

Example 1
Summary: Speaker_1 is sharing details about their recent trip to India with Speaker_2. The conversation reveals that Speaker_1 had a positive experience with Indian food during their visit, noting that restaurants catered to Americans by making dishes less spicy, and this experience has led them to frequent a local Indian restaurant weekly since returning home. Speaker_1 describes visiting a bird sanctuary in southern India as part of their trip, which involved a group of about 10 people. While some members of their group were able to extend their stay to visit additional attractions like an elephant reserve, Speaker_1 had to return due to work and class commitments they couldn't miss. The conversation also touches on the weather conditions during their visit, with Speaker_1 mentioning that it wasn't too hot during their time there, though their host indicated they had just experienced a particularly intense hot season. Throughout the dialogue, Speaker_2 appears to be asking follow-up questions about various aspects of the trip, showing interest in Speaker_1's travel experiences in India.
Content Coverage: 4
Covers main points but lacks detail on Speaker_2's contributions.
Dialogue Flow: 4
Maintains logical flow but less emphasis on Speaker_2's role.
Information Accuracy: 4
Accurate but omits some context from Speaker_2.
Purpose Outcome 4
Conveys trip purpose but less on dialogue outcome.
Detail Balance: 3
Focuses more on Speaker_1, less on Speaker_2.
Summary: The conversation centers around Speaker_1 sharing their recent travel experience to India with Speaker_2, who shows interest and asks follow-up questions about the trip. Speaker_1 describes enjoying the local Indian food, which was prepared in a way that catered to Americans by not being too spicy, and mentions that this experience led them to discover a local Indian restaurant they now visit weekly. The discussion moves to other activities during the trip, including a visit to a bird sanctuary in southern India, though Speaker_1 doesn't recall specific names or provide details about the birds they saw. Speaker_1 traveled with a group of about 10 people, some of whom were able to extend their stay to visit additional attractions like an elephant reserve, but Speaker_1 had to return due to work and class commitments. The conversation also touches on the weather conditions during the visit, with Speaker_1 noting that while it wasn't too hot during their stay, their host mentioned it was just before a particularly intense hot season, making their timing fortunate. Throughout the exchange, Speaker_2 maintains an encouraging and interested tone, expressing that the trip sounds like an amazing educational experience despite being relatively short.
Content Coverage: 4
Covers key points but lacks depth in some areas.
Dialogue Flow: 4
Flow is maintained but less emphasis on interaction.
Information Accuracy: 4
Accurate but lacks some specific details.
Purpose Outcome 4
Conveys purpose but less on specific outcomes.
Detail Balance: 4
Balances details but lacks depth in Speaker_2's input.
Summary: This conversation centers around Speaker_1's recent trip to India, likely as part of a university program. Speaker_1 discusses their positive experience with Indian food, noting that it was prepared to accommodate American tastes and wasn't too spicy, which has led them to regularly visit an Indian restaurant near their home since returning. The trip included visits to various cities in southern India, including a bird sanctuary, with Speaker_2 explaining that the southern region offers more tropical environments and natural reserves. Speaker_1 traveled with about 10 other students from their university, though some participants were able to extend their stay to experience the Holi color festival and visit an elephant reserve, which Speaker_1 couldn't do due to work and class commitments back home. The timing of the trip was favorable as it occurred before India's intense hot season begins in mid-March, with Speaker_2 noting that February or March is an ideal time to visit when the weather is more pleasant and major festivals like Holi take place.
Content Coverage: 5
Covers all key points including food, travel, group size, and timing.
Dialogue Flow: 5
Reflects natural progression with clear transitions between topics.
Information Accuracy: 5
Accurately represents dialogue details and context.
Purpose Outcome 5
Clearly conveys trip goals and outcomes.
Detail Balance: 5
Balances details from both speakers effectively.
Example 2
Summary: The conversation centers around Speaker_1 discovering that a B-movie they watched on Amazon was actually filmed at their school, UC Riverside. Speaker_1 explains they were watching a collection of bad movies presented by Elvira called \"Movie Macabre\" when they recognized familiar locations from their campus in one of the films, which appeared to be titled something like \"Attack of the Amazon\" or similar. Initially, Speaker_1 didn't immediately recognize the setting since the movie was shot in the 1980s and the campus has many buildings, but when the distinctive bell tower appeared on screen, they realized it was their university. After googling the film, Speaker_1 confirmed that the movie was indeed shot at UC Riverside, and they were able to identify other campus locations like the library that appeared in the film. The conversation appears to shift toward the end when Speaker_2 brings up something about snow and asks about someone knowing what movie it was, though Speaker_1 seems uncertain about this new topic.
Content Coverage: 3
Focuses mainly on Speaker_1's story, missing Speaker_2's film crew experience.
Dialogue Flow: 3
Dialogue flow is somewhat abrupt, missing smooth transition.
Information Accuracy: 3
Accurate but incomplete, missing key parts of Speaker_2's story.
Purpose Outcome 3
Partially conveys the dialogue's goals, missing full outcome.
Detail Balance: 2
Heavily weighted towards Speaker_1's perspective.
Summary: The conversation begins with the speakers discussing character development in media, specifically how compelling characters can evoke empathy even when they engage in questionable behavior. Speaker_1 then shifts the topic to describe discovering a collection of B-movies hosted by Elvira on Amazon, including one called \"Attack of the Amazon\" or something similar. Speaker_1 becomes excited when recognizing that this 1980s film was shot at their own school campus, UC Riverside, particularly noting the distinctive bell tower and library that helped them identify the location. Speaker_2 expresses interest in this discovery and comments on how surreal it must be to see familiar places in a decades-old movie context, suggesting the campus was likely chosen for its classic college appearance. The conversation then takes an abrupt turn when Speaker_2 begins describing the animated series \"Disenchantment\" by Matt Groening, detailing the show's medieval fantasy setting, main character Princess Bean, and distinctive animation style. The dialogue concludes with a brief, somewhat unclear exchange about someone's knowledge of a movie, possibly referencing Snow White and the Seven Dwarfs, though the connection to the previous topics remains unclear.
Content Coverage: 2
Introduces unrelated topics, missing key parts of the dialogue.
Dialogue Flow: 2
Abrupt topic shifts disrupt natural dialogue flow.
Information Accuracy: 2
Includes inaccuracies and unrelated content.
Purpose Outcome 2
Fails to clearly convey the dialogue's goals and outcomes.
Detail Balance: 2
Imbalanced with irrelevant details overshadowing main points.
Summary: The conversation centers around the speakers sharing experiences with B-movies and film productions. Speaker_1 begins by discussing watching old horror movies from the 50s and 60s, mentioning a collection called \"Movie Macabre\" hosted by Elvira on Amazon that featured bad movies. They excitedly recount discovering that one of these films was actually shot at their school, UC Riverside, recognizing the distinctive bell tower and library in the movie. Speaker_2 then shares a related story about a woman they used to rent from who owned a Victorian house that she had listed online as a filming location. The woman expected small commercial shoots but instead ended up hosting a horror film crew for two to three months, which turned into a nightmare experience with constant screaming, disruptive sounds, and a massive cleanup job that required professional services afterward. The woman told Speaker_2 she wouldn't rent her house for filming again, though Speaker_2 acknowledged it made for an interesting story. The conversation flows naturally from discussing watching B-movies to personal connections with film production, highlighting how filming can unexpectedly intersect with everyday life.
Content Coverage: 5
Covers both main stories: B-movie discovery and horror film crew experience.
Dialogue Flow: 5
Reflects natural progression from B-movies to personal film experiences.
Information Accuracy: 5
Accurately captures key details and interactions.
Purpose Outcome 5
Clearly conveys the dialogue's goals and outcomes.
Detail Balance: 5
Balances details from both speakers effectively.
Example 3
This conversation appears to be a casual discussion between two people covering multiple topics. The dialogue begins with Speaker_1 discussing psychology and memory research, explaining the current limitations and unknowns in the field. Speaker_1, who identifies as a memory researcher, discusses theoretical questions about where memories are stored in the brain - whether in individual neural cells or their connections - and draws an analogy to computer storage, suggesting there must be capacity limits since we can't remember every experience in perfect detail. The conversation then shifts to discussing weather and location, with Speaker_1 describing the pleasant climate in Santa Barbara, noting temperatures in the 70s with recent fog, and explaining how the area maintains moderate temperatures year-round, rarely exceeding 90 degrees in summer or dropping below 40-50 degrees in winter. Speaker_1 expresses a preference for Southern California over Northern California weather. The discussion concludes with references to a particularly brutal summer that Speaker_2 apparently experienced, though the specific details of Speaker_2's contributions remain unclear due to the masked responses.
Content Coverage: 3
Misses some details about animal instincts and smoke conditions.
Dialogue Flow: 4
Maintains logical flow but lacks some transitions.
Information Accuracy: 4
Accurate but omits some speaker contributions.
Purpose Outcome 4
Conveys main goals but lacks some outcomes.
Detail Balance: 3
Focuses more on Speaker_1's contributions.
This conversation involves two speakers discussing the fascinating complexities of human memory and brain function, followed by a shift to discussing weather conditions. The speakers explore how the brain optimizes memory storage through cost-benefit analysis, determining what information is worth storing versus what can be reconstructed later. Speaker_1, who identifies as a memory researcher, acknowledges that despite their expertise, there remains much unknown about psychology and memory storage, including fundamental questions about where memories are actually stored - whether in individual neural cells or in their connections - and the brain's capacity limits. They discuss how the brain must make trade-offs between storage capacity and retrieval efficiency, noting that storing everything in perfect detail would likely overwhelm our ability to access relevant memories when needed. The conversation then transitions to weather, with Speaker_1 describing the pleasant climate in Santa Barbara, where temperatures stay in the 70s with mild seasonal variations, contrasting this with Speaker_2's experience of recent extreme heat reaching the hundreds for several days, which strained the power grid and was compounded by wildfires and smoke. The speakers agree that Santa Barbara's consistent, moderate weather is preferable to the extreme conditions experienced elsewhere.
Content Coverage: 4
Includes memory and weather but adds inaccurate details about power grid.
Dialogue Flow: 3
Flow is logical but introduces non-existent topics.
Information Accuracy: 3
Inaccurate details about extreme heat and power grid.
Purpose Outcome 3
Conveys goals but introduces incorrect outcomes.
Detail Balance: 3
Balances details but includes inaccuracies.
This conversation involves two speakers discussing the mysteries of memory storage and brain function, with Speaker_1 being a memory researcher. The discussion begins with comparisons between computer data storage and how the human brain stores memories, with Speaker_2 expressing fascination about the seemingly magical nature of brain function despite having taken biology classes. Speaker_1 acknowledges that even as a memory researcher, much about memory storage remains unknown, including whether memories are stored in individual neural cells or in their connections, and notes that like computers, brains must have storage capacity limits. The conversation expands to discuss animal instincts and where that information might be stored genetically. The dialogue then shifts to a more casual discussion about weather conditions, with Speaker_1 describing the pleasant climate in Santa Barbara, noting temperatures in the 70s with some recent fog, and comparing it favorably to San Diego's climate. Speaker_2 mentions disliking heat and describes difficult conditions in their location during summer, including both extreme heat and poor air quality from smoke that made it undesirable to go outside, compounding the isolation already experienced due to Covid restrictions.
Content Coverage: 5
Covers key topics: memory, brain function, weather, and smoke conditions
Dialogue Flow: 5
Reflects natural progression from memory to weather discussion.
Information Accuracy: 5
Accurately represents dialogue content and speaker roles.
Purpose Outcome 5
Clearly conveys dialogue goals and outcomes.
Detail Balance: 5
Balances details from both speakers effectively.

Insights

  • ℹ️
    Access to more information improves reconstruction In a one-sided environment, information about the length of a masked turn and access to the turn immediately following a masked turn improves reconstruction quality.
  • 🤖
    Plausible reconstructions are possible without fine-tuning Large pretrained models can generate plausible reconstructions out-of-the-box; even with finetuning, smaller models do not achieve the same capabilities.
  • 📋
    High-quality summaries can be produced directly from one-sided input Summary generation can be performed on one-sided input without reconstructing the missing turns. Furthermore, while models struggle with reconstructing missing turns from speech-based transcriptions, the summaries produced from these conversations are on par with those of text-based conversations.

Results

Table 1. Effect of context ablations on reconstruction quality — mean (SD), scored 1–5 by GPT-4o. Overall best performing set up is summary-highlighted in blue.
Dataset N+1 Turn Len. "xxxx" Instr. Full Prior Context Prompted / Finetuned Claude / Llama Seman. Sim. (↑) Intent Pres. (↑) Context. Approp. (↑) Summ. Align. (↑) Anti-Halluc. (↑)
DailyDialog PC 1.79 (0.95)2.65 (1.34)3.11 (1.09)1.80 (0.96)4.05 (1.61)
PC 1.85 (0.92)2.75 (1.37)3.32 (1.08)1.89 (0.94)4.60 (0.99)
PC 2.34 (1.07)3.42 (1.32)3.60 (1.06)2.39 (1.13)4.65 (0.92)
PC 2.14 (1.18)2.97 (1.49)3.77 (1.10)2.81 (1.29)4.66 (0.90)
PC 2.63 (1.28)3.60 (1.36)3.80 (1.05)2.71 (1.31)4.77 (0.76)
PC 2.12 (1.11)2.95 (1.46)3.44 (1.12)2.17 (1.15)4.60 (0.96)
PL 1.21 (0.47)1.41 (0.76)1.57 (0.85)1.20 (0.45)1.87 (1.16)
FL 1.38 (0.88)1.80 (1.27)1.64 (1.03)1.39 (0.90)2.49 (1.60)
MultiWOZ PC 2.50 (1.09)3.54 (1.38)3.39 (1.19)2.44 (1.09)3.14 (1.91)
PC 2.50 (1.01)3.60 (1.33)3.69 (1.08)2.50 (1.04)4.71 (0.77)
PC 2.64 (1.00)3.87 (1.21)3.76 (1.01)2.66 (1.03)4.59 (0.92)
PC 2.84 (1.27)3.81 (1.36)3.95 (1.12)2.81 (1.29)3.98 (1.65)
PC 2.96 (1.20)4.06 (1.17)4.07 (0.98)3.00 (1.24)4.74 (0.75)
PC 2.59 (1.06)3.70 (1.29)3.76 (1.05)2.61 (1.09)4.67 (0.82)
PL 1.63 (0.66)1.98 (0.95)2.06 (0.90)1.60 (0.61)2.24 (1.13)
FL 1.98 (1.12)3.13 (1.42)2.39 (1.17)1.98 (1.11)2.46 (1.46)
SpokenWOZ PC 1.62 (0.99)2.19 (1.39)2.13 (1.24)1.60 (0.98)1.91 (1.18)
PC 1.67 (0.99)2.22 (1.38)2.22 (1.24)1.66 (1.01)2.18 (1.32)
PC 1.85 (1.11)2.51 (1.47)2.30 (1.27)1.85 (1.11)2.28 (1.33)
PC 2.07 (1.32)2.65 (1.61)2.67 (1.42)2.11 (1.37)3.14 (1.54)
PC 2.29 (1.36)2.96 (1.60)2.76 (1.36)2.31 (1.40)3.22 (1.52)
Candor PC 1.18 (0.48)1.48 (0.82)1.86 (0.92)1.18 (0.46)1.99 (1.20)
PC 1.19 (0.50)1.53 (0.85)1.94 (0.95)1.19 (0.49)2.17 (1.31)
PC 1.24 (0.58)1.58 (0.97)1.82 (0.94)1.23 (0.57)2.13 (1.25)
PC 1.62 (1.10)2.04 (1.32)2.48 (1.22)1.69 (1.20)3.87 (1.46)
PC 1.61 (1.10)2.06 (1.32)2.43 (1.22)1.68 (1.22)3.79 (1.28)
Table 2. Effect of context masking on GPT-4o-judged summary quality — mean (SD), scored 1–5. Overall best performing set up is summary-highlighted in blue.
Dataset Condition Content Coverage Information Accuracy Dialogue Flow Purpose/Outcome Detail Balance
DailyDialog Full 4.83 (0.41) 4.87 (0.41) 4.80 (0.45) 4.89 (0.36) 4.84 (0.42)
Masked 3.44 (0.64) 3.51 (0.69) 3.58 (0.68) 3.59 (0.79) 3.38 (0.70)
Predicted 3.04 (0.94) 2.61 (1.08) 3.11 (0.93) 3.09 (1.04) 2.99 (0.97)
MultiWOZ Full 4.95 (0.22) 4.95 (0.24) 4.91 (0.29) 4.98 (0.15) 4.97 (0.18)
Masked 3.35 (0.53) 3.40 (0.55) 3.69 (0.62) 3.74 (0.67) 3.41 (0.57)
Predicted 3.26 (0.74) 3.01 (0.89) 3.54 (0.73) 3.56 (0.83) 3.28 (0.80)
SpokenWOZ Full 4.61 (0.62) 4.57 (0.67) 4.54 (0.59) 4.81 (0.42) 4.65 (0.61)
Masked 3.68 (0.63) 3.60 (0.70) 3.80 (0.71) 4.03 (0.81) 3.75 (0.74)
Predicted 3.17 (0.71) 2.74 (0.86) 3.27 (0.78) 3.37 (0.88) 3.14 (0.84)
Candor (avg. 426 turns) Full 4.60 (0.52) 4.60 (0.52) 4.60 (0.52) 4.60 (0.52) 4.60 (0.52)
Masked 3.40 (0.52) 3.50 (0.53) 3.40 (0.52) 3.40 (0.52) 3.10 (0.44)
Predicted 4.00 (0.82) 3.90 (0.99) 4.10 (0.74) 4.20 (0.79) 4.00 (0.94)
Short Candor (avg. 25 turns) Full 4.65 (0.51) 4.78 (0.47) 4.50 (0.53) 4.74 (0.47) 4.66 (0.52)
Masked 3.01 (0.80) 3.15 (0.87) 3.25 (0.67) 3.09 (0.85) 2.80 (0.78)
Predicted 2.72 (0.95) 2.47 (0.47) 2.81 (0.91) 2.77 (1.08) 2.63 (0.96)