Skip to content

dotnet aieval report Cases tab: "No transcript for this case" gives no indication the transcript was intentionally stripped for non-latest executions #7759

Description

@jzuras

dotnet aieval report intentionally clears ScenarioRunResult.Messages/ModelResponse for every execution except the most recent one returned by GetLatestExecutionNamesAsync (see ReportCommand.InvokeAsync, the else branch with the comment // Clear the chat data for following executions). This is a reasonable size-control tradeoff for the report file. However, the Cases tab is scoped to an individual run's results, selectable with a dropdown to switch between executions of the same scenario/iteration, and it renders this as:

"No transcript for this case."

That message is indistinguishable from an actual data problem: a missing/broken transcript, a bug in the evaluation code, or a serialization failure. Nothing in the UI discloses that this is by design or that the data was ever present. This sent me on a multi-day investigation after creating my own multi-judge evaluation unit test (comparing on-disk ScenarioRunResult JSON against the report's embedded dataset, inspecting TranscriptBlock.tsx, etc.) before finding the actual cause in ReportCommand.cs, all of which would have been unnecessary with a more specific message.

Suggested fixes (either of these would help):

Change the message when the underlying cause is known stripping, e.g.: "Transcript not retained for this execution (only the most recent execution's transcript is kept in the report)." Even a generic version ("Transcript omitted for older executions to keep report size manageable") would have immediately pointed me at the right explanation instead of "the data must be broken somewhere."
Add a CLI option to control retention, e.g. --keep-transcripts or --no-strip-transcripts, so users doing small-scale comparisons (a handful of executions — different judges, models, or prompt variants over the same scenario) can opt out of the stripping instead of losing all but one transcript regardless of how many executions are actually in the report. This may be a low-priority ask relative to the messaging fix, but worth mentioning since the current behavior applies uniformly regardless of how many executions are actually being reported (i.e. it strips just as aggressively for 2 executions as for 200).

Repro: Using the official sample setup (DiskBasedReportingConfiguration.Create(...), as shown in the Microsoft.Extensions.AI.Evaluation walkthroughs), run the same test (same scenario/iteration) twice so two distinct executions are persisted, then generate the report with dotnet aieval report. Open the Cases tab and switch the execution dropdown to the older run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-ai-evalMicrosoft.Extensions.AI.Evaluation and relateduntriaged

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions