Discovering capabilities under real world constraints.
Conversational AI is constrained in many real-world settings where only one side of a dialogue can be recorded. We formalize the one-sided conversation problem (1SC): inferring and learning from only one side of a conversation. We study two tasks: (1) reconstructing the missing speaker's turns and (2) generating summaries from one-sided transcripts. Evaluating models on MultiWOZ, DailyDialog, SpokenWOZ and Candor with both human A/B testing and LLM-as-a-judge metrics, we find that additional context improves reconstruction, and while large models generate promising reconstructions with prompting, smaller models require finetuning. Further, high-quality summaries can be generated without reconstructing missing turns. We present 1SC as a novel challenge and report promising results that mark a step toward privacy-aware conversational AI.
| Dataset | N+1 | Turn Len. | "xxxx" Instr. | Full Prior Context | Prompted / Finetuned | Claude / Llama | Seman. Sim. (↑) | Intent Pres. (↑) | Context. Approp. (↑) | Summ. Align. (↑) | Anti-Halluc. (↑) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DailyDialog | ✗ | ✗ | ✗ | ✓ | P | C | 1.79 (0.95) | 2.65 (1.34) | 3.11 (1.09) | 1.80 (0.96) | 4.05 (1.61) |
| ✗ | ✗ | ✓ | ✓ | P | C | 1.85 (0.92) | 2.75 (1.37) | 3.32 (1.08) | 1.89 (0.94) | 4.60 (0.99) | |
| ✓ | ✗ | ✓ | ✓ | P | C | 2.34 (1.07) | 3.42 (1.32) | 3.60 (1.06) | 2.39 (1.13) | 4.65 (0.92) | |
| ✗ | ✓ | ✓ | ✓ | P | C | 2.14 (1.18) | 2.97 (1.49) | 3.77 (1.10) | 2.81 (1.29) | 4.66 (0.90) | |
| ✓ | ✓ | ✓ | ✓ | P | C | 2.63 (1.28) | 3.60 (1.36) | 3.80 (1.05) | 2.71 (1.31) | 4.77 (0.76) | |
| ✓ | ✗ | ✓ | ✗ | P | C | 2.12 (1.11) | 2.95 (1.46) | 3.44 (1.12) | 2.17 (1.15) | 4.60 (0.96) | |
| ✓ | ✗ | ✓ | ✗ | P | L | 1.21 (0.47) | 1.41 (0.76) | 1.57 (0.85) | 1.20 (0.45) | 1.87 (1.16) | |
| ✓ | ✗ | ✓ | ✗ | F | L | 1.38 (0.88) | 1.80 (1.27) | 1.64 (1.03) | 1.39 (0.90) | 2.49 (1.60) | |
| MultiWOZ | ✗ | ✗ | ✗ | ✓ | P | C | 2.50 (1.09) | 3.54 (1.38) | 3.39 (1.19) | 2.44 (1.09) | 3.14 (1.91) |
| ✗ | ✗ | ✓ | ✓ | P | C | 2.50 (1.01) | 3.60 (1.33) | 3.69 (1.08) | 2.50 (1.04) | 4.71 (0.77) | |
| ✓ | ✗ | ✓ | ✓ | P | C | 2.64 (1.00) | 3.87 (1.21) | 3.76 (1.01) | 2.66 (1.03) | 4.59 (0.92) | |
| ✗ | ✓ | ✓ | ✓ | P | C | 2.84 (1.27) | 3.81 (1.36) | 3.95 (1.12) | 2.81 (1.29) | 3.98 (1.65) | |
| ✓ | ✓ | ✓ | ✓ | P | C | 2.96 (1.20) | 4.06 (1.17) | 4.07 (0.98) | 3.00 (1.24) | 4.74 (0.75) | |
| ✓ | ✗ | ✓ | ✗ | P | C | 2.59 (1.06) | 3.70 (1.29) | 3.76 (1.05) | 2.61 (1.09) | 4.67 (0.82) | |
| ✓ | ✗ | ✓ | ✗ | P | L | 1.63 (0.66) | 1.98 (0.95) | 2.06 (0.90) | 1.60 (0.61) | 2.24 (1.13) | |
| ✓ | ✗ | ✓ | ✗ | F | L | 1.98 (1.12) | 3.13 (1.42) | 2.39 (1.17) | 1.98 (1.11) | 2.46 (1.46) | |
| SpokenWOZ | ✗ | ✗ | ✗ | ✓ | P | C | 1.62 (0.99) | 2.19 (1.39) | 2.13 (1.24) | 1.60 (0.98) | 1.91 (1.18) |
| ✗ | ✗ | ✓ | ✓ | P | C | 1.67 (0.99) | 2.22 (1.38) | 2.22 (1.24) | 1.66 (1.01) | 2.18 (1.32) | |
| ✓ | ✗ | ✓ | ✓ | P | C | 1.85 (1.11) | 2.51 (1.47) | 2.30 (1.27) | 1.85 (1.11) | 2.28 (1.33) | |
| ✗ | ✓ | ✓ | ✓ | P | C | 2.07 (1.32) | 2.65 (1.61) | 2.67 (1.42) | 2.11 (1.37) | 3.14 (1.54) | |
| ✓ | ✓ | ✓ | ✓ | P | C | 2.29 (1.36) | 2.96 (1.60) | 2.76 (1.36) | 2.31 (1.40) | 3.22 (1.52) | |
| Candor | ✗ | ✗ | ✗ | ✓ | P | C | 1.18 (0.48) | 1.48 (0.82) | 1.86 (0.92) | 1.18 (0.46) | 1.99 (1.20) |
| ✗ | ✗ | ✓ | ✓ | P | C | 1.19 (0.50) | 1.53 (0.85) | 1.94 (0.95) | 1.19 (0.49) | 2.17 (1.31) | |
| ✓ | ✗ | ✓ | ✓ | P | C | 1.24 (0.58) | 1.58 (0.97) | 1.82 (0.94) | 1.23 (0.57) | 2.13 (1.25) | |
| ✗ | ✓ | ✓ | ✓ | P | C | 1.62 (1.10) | 2.04 (1.32) | 2.48 (1.22) | 1.69 (1.20) | 3.87 (1.46) | |
| ✓ | ✓ | ✓ | ✓ | P | C | 1.61 (1.10) | 2.06 (1.32) | 2.43 (1.22) | 1.68 (1.22) | 3.79 (1.28) |
| Dataset | Condition | Content Coverage | Information Accuracy | Dialogue Flow | Purpose/Outcome | Detail Balance |
|---|---|---|---|---|---|---|
| DailyDialog | Full | 4.83 (0.41) | 4.87 (0.41) | 4.80 (0.45) | 4.89 (0.36) | 4.84 (0.42) |
| Masked | 3.44 (0.64) | 3.51 (0.69) | 3.58 (0.68) | 3.59 (0.79) | 3.38 (0.70) | |
| Predicted | 3.04 (0.94) | 2.61 (1.08) | 3.11 (0.93) | 3.09 (1.04) | 2.99 (0.97) | |
| MultiWOZ | Full | 4.95 (0.22) | 4.95 (0.24) | 4.91 (0.29) | 4.98 (0.15) | 4.97 (0.18) |
| Masked | 3.35 (0.53) | 3.40 (0.55) | 3.69 (0.62) | 3.74 (0.67) | 3.41 (0.57) | |
| Predicted | 3.26 (0.74) | 3.01 (0.89) | 3.54 (0.73) | 3.56 (0.83) | 3.28 (0.80) | |
| SpokenWOZ | Full | 4.61 (0.62) | 4.57 (0.67) | 4.54 (0.59) | 4.81 (0.42) | 4.65 (0.61) |
| Masked | 3.68 (0.63) | 3.60 (0.70) | 3.80 (0.71) | 4.03 (0.81) | 3.75 (0.74) | |
| Predicted | 3.17 (0.71) | 2.74 (0.86) | 3.27 (0.78) | 3.37 (0.88) | 3.14 (0.84) | |
| Candor (avg. 426 turns) | Full | 4.60 (0.52) | 4.60 (0.52) | 4.60 (0.52) | 4.60 (0.52) | 4.60 (0.52) |
| Masked | 3.40 (0.52) | 3.50 (0.53) | 3.40 (0.52) | 3.40 (0.52) | 3.10 (0.44) | |
| Predicted | 4.00 (0.82) | 3.90 (0.99) | 4.10 (0.74) | 4.20 (0.79) | 4.00 (0.94) | |
| Short Candor (avg. 25 turns) | Full | 4.65 (0.51) | 4.78 (0.47) | 4.50 (0.53) | 4.74 (0.47) | 4.66 (0.52) |
| Masked | 3.01 (0.80) | 3.15 (0.87) | 3.25 (0.67) | 3.09 (0.85) | 2.80 (0.78) | |
| Predicted | 2.72 (0.95) | 2.47 (0.47) | 2.81 (0.91) | 2.77 (1.08) | 2.63 (0.96) |