dotnet aieval report intentionally clears ScenarioRunResult.Messages/ModelResponse for every execution except the most recent one returned by GetLatestExecutionNamesAsync (see ReportCommand.InvokeAsync, the else branch with the comment // Clear the chat data for following executions). This is a reasonable size-control tradeoff for the report file. However, the Cases tab is scoped to an individual run's results, selectable with a dropdown to switch between executions of the same scenario/iteration, and it renders this as:
"No transcript for this case."
That message is indistinguishable from an actual data problem: a missing/broken transcript, a bug in the evaluation code, or a serialization failure. Nothing in the UI discloses that this is by design or that the data was ever present. This sent me on a multi-day investigation after creating my own multi-judge evaluation unit test (comparing on-disk ScenarioRunResult JSON against the report's embedded dataset, inspecting TranscriptBlock.tsx, etc.) before finding the actual cause in ReportCommand.cs, all of which would have been unnecessary with a more specific message.
Suggested fixes (either of these would help):
Change the message when the underlying cause is known stripping, e.g.: "Transcript not retained for this execution (only the most recent execution's transcript is kept in the report)." Even a generic version ("Transcript omitted for older executions to keep report size manageable") would have immediately pointed me at the right explanation instead of "the data must be broken somewhere."
Add a CLI option to control retention, e.g. --keep-transcripts or --no-strip-transcripts, so users doing small-scale comparisons (a handful of executions — different judges, models, or prompt variants over the same scenario) can opt out of the stripping instead of losing all but one transcript regardless of how many executions are actually in the report. This may be a low-priority ask relative to the messaging fix, but worth mentioning since the current behavior applies uniformly regardless of how many executions are actually being reported (i.e. it strips just as aggressively for 2 executions as for 200).
Repro: Using the official sample setup (DiskBasedReportingConfiguration.Create(...), as shown in the Microsoft.Extensions.AI.Evaluation walkthroughs), run the same test (same scenario/iteration) twice so two distinct executions are persisted, then generate the report with dotnet aieval report. Open the Cases tab and switch the execution dropdown to the older run.
dotnet aieval report intentionally clears ScenarioRunResult.Messages/ModelResponse for every execution except the most recent one returned by GetLatestExecutionNamesAsync (see ReportCommand.InvokeAsync, the else branch with the comment // Clear the chat data for following executions). This is a reasonable size-control tradeoff for the report file. However, the Cases tab is scoped to an individual run's results, selectable with a dropdown to switch between executions of the same scenario/iteration, and it renders this as:
"No transcript for this case."
That message is indistinguishable from an actual data problem: a missing/broken transcript, a bug in the evaluation code, or a serialization failure. Nothing in the UI discloses that this is by design or that the data was ever present. This sent me on a multi-day investigation after creating my own multi-judge evaluation unit test (comparing on-disk ScenarioRunResult JSON against the report's embedded dataset, inspecting TranscriptBlock.tsx, etc.) before finding the actual cause in ReportCommand.cs, all of which would have been unnecessary with a more specific message.
Suggested fixes (either of these would help):
Change the message when the underlying cause is known stripping, e.g.: "Transcript not retained for this execution (only the most recent execution's transcript is kept in the report)." Even a generic version ("Transcript omitted for older executions to keep report size manageable") would have immediately pointed me at the right explanation instead of "the data must be broken somewhere."
Add a CLI option to control retention, e.g. --keep-transcripts or --no-strip-transcripts, so users doing small-scale comparisons (a handful of executions — different judges, models, or prompt variants over the same scenario) can opt out of the stripping instead of losing all but one transcript regardless of how many executions are actually in the report. This may be a low-priority ask relative to the messaging fix, but worth mentioning since the current behavior applies uniformly regardless of how many executions are actually being reported (i.e. it strips just as aggressively for 2 executions as for 200).
Repro: Using the official sample setup (DiskBasedReportingConfiguration.Create(...), as shown in the Microsoft.Extensions.AI.Evaluation walkthroughs), run the same test (same scenario/iteration) twice so two distinct executions are persisted, then generate the report with dotnet aieval report. Open the Cases tab and switch the execution dropdown to the older run.