Here's the idea: build an app that checks if someone knows their stuff using an interview. You pick a skill, type your name, and off you go. An AI agent chats with you by voice, asks questions, and at the end you get a report card. Hopefully with more compliments than trauma.
I've seen some applications that evaluate interview participants with AI, so I decided to build one out of curiosity. It was an interesting journey. The trickiest part was calibrating the prompt, especially to end the interview at the right time.
Let's get the requirements first, so we have a clear picture of what to implement, before I start the architecture, the code, drawing boxes and arrows, and pretending I know what I'm doing lol.
Every skill is a JSON file. Adding a new skill means adding a file: no code changes. A skill defines:
| Field | Purpose |
|---|---|
Title |
Name shown to the participant, e.g. "SQL" |
InterviewInstruction |
What the agent should ask and which topics matter most |
AgentName ,AgentTone |
Who the agent is and how it behaves (friendly, professional, like a Jedi master...) |
AgentVoiceGender ,AgentVoiceType |
The agent's voice (Male/Female; HighPitched, Neutral, Deep, Warm) |
Effort |
How capable the AI model should be: Low, Medium or High |
ReportInstruction |
How to evaluate and what to put in the report |
MaxPoints |
The highest possible score |
InterviewDurationInMinutes ,ExtraInterviewDurationInMinutes |
Planned duration and extra time |
PointInstructionMap |
A label and the minimum requirements for every point from 1 toMaxPoints |
For example, a Star Wars skill can use fun level names:
{
"Id": "star-wars",
"Title": "Star Wars Lore",
"InterviewInstruction": "Test the participant's knowledge of the Star Wars universe...",
"AgentTone": "Warm and playful, like a wise old Jedi master.",
"AgentName": "Master Oren",
"AgentVoiceGender": "Male",
"AgentVoiceType": "Deep",
"Effort": "Low",
"ReportInstruction": "Evaluate breadth and depth of lore knowledge.",
"MaxPoints": 4,
"InterviewDurationInMinutes": 10,
"ExtraInterviewDurationInMinutes": 2,
"PointInstructionMap": {
"1": { "Label": "Youngling", "Requirements": "Knows the main characters." },
"2": { "Label": "Padawan", "Requirements": "Knows the plot of the main films." },
"3": { "Label": "Jedi Knight", "Requirements": "Knows the history of the Jedi and the Sith." },
"4": { "Label": "Jedi Master", "Requirements": "Knows deep lore, including series and books." }
}
}
If a skill file is invalid (for example, a missing point in PointInstructionMap), the app starts with a clear message saying what to fix.
This is a sample for an article, so some things are deliberately left out:
The interview logic is agnostic and doesn't depend on a specific AI vendor.
The interviewer and the report agent (OpenAI) and the speech-to-text/text-to-speech service (** Azure Speech**) are each behind an interface, in their own project, so that they can be replaced by different providers.
Let's build the abstractions first: the backend components, and how the interview engine handles everything the participant says so the conversation flows as naturally as possible.
The engine is split into two parts:
The participant's voice is transcribed and sent to the agent; then the agent writes its reply, and the reply is synthesized into speech with the voice from the config.
Every box in the bottom row is an interface, implemented in its own project (SkillsValidator.Engine.OpenAI, SkillsValidator.Engine.AzureSpeech). The engine never references a vendor SDK.
Skills are read-only: they come from the JSON files.
public interface ISkillConfigRepository
{
Task<IReadOnlyList<SkillConfig>> GetAllAsync();
Task<SkillConfig?> GetAsync(string configId);
SkillConfigLoadResult GetLoadResult(); // startup validation: the errors to show if a file is invalid
}
An interview is a session. It goes through three states: Running → Finished (evaluation pending) → Reported (the report is ready).
public interface ISessionInterviewService
{
Task<string> StartInterviewSessionAsync(string participantName, SkillConfig config);
Task FinishInterviewSessionAsync(string sessionId); // also starts the evaluation
Task<InterviewSession?> GetSessionAsync(string sessionId);
}
GetSessionAsync returns the whole picture: the participant, the state, the transcript and, once Reported, the result. Behind it, three small repositories (session, transcript, result) keep everything in memory. The transcript is append-only: an interview only ever adds new lines.
When the interview finishes, a second agent evaluates it:
public interface IReportEvaluator
{
Task<SessionReportResult> EvaluateAsync(
SkillConfig config, IReadOnlyList<TranscriptEntry> transcript, CancellationToken ct = default);
}
Four interfaces make the voice conversation work.
1. Hearing and speaking: ISpeechConverter turns audio into text and text into audio.
public interface ISpeechConverter
{
// Microphone audio in, text out: partial text while the participant speaks, final text after a short silence.
IAsyncEnumerable<TranscriptionSegment> ToTranscriptionAsync(
IAsyncEnumerable<AudioChunk> speech, CancellationToken ct = default);
// Text in, the agent's voice out.
IAsyncEnumerable<AudioChunk> ToSpeechAsync(
string text, VoiceGender gender, VoiceType type, string? speakingStyle, CancellationToken ct = default);
}
public sealed record AudioChunk(byte[] Data, string Format, int SampleRate);
public sealed record TranscriptionSegment(string Text, bool IsFinal);
Notice that everything is an IAsyncEnumerable: audio and text flow through the system in small pieces. The engine never waits for a whole answer before doing the next step.
2. Thinking: IRealtimeInterviewAgent is the interviewer. Given the skill and the conversation so far, it streams its next reply as text fragments.
public interface IRealtimeInterviewAgent
{
IAsyncEnumerable<string> RespondAsync(SkillConfig config, InterviewSession session, CancellationToken ct = default);
}
3. The pipe to the browser: IAudioChannel receives the participant's microphone audio and plays the agent's voice.
public interface IAudioChannel
{
IAsyncEnumerable<AudioChunk> ReadParticipantAudioAsync(CancellationToken ct);
Task PlayAgentAudioAsync(IAsyncEnumerable<AudioChunk> audio, CancellationToken ct);
Task StopAgentAudioAsync(); // the participant interrupted
}
The engine doesn't care how the audio travels. In our app it's a WebSocket.
4. The conductor: IRealtimeInterviewRunner connects the other three and runs the conversation until it ends.
public interface IRealtimeInterviewRunner
{
Task RunAsync(string sessionId, SkillConfig config, IAudioChannel channel,
Func<string, Task>? onPartialTranscript, // live caption
CancellationToken ct);
}
The runner listens all the time. The speech converter sends it two kinds of events, and each one gets a different reaction:
While the participant is speaking (a partial segment):
StopAgentAudioAsync silences the speaker. This is what makes When the participant has finished (a final segment, after about 2 seconds of silence):
Speaking sentence by sentence is the key to a natural flow. The participant hears the first sentence after roughly a second, instead of waiting for the whole reply to be written and then spoken.
Keeping the agent on track. Two details make the conversation feel less robotic:
Ending. When the agent decides to close the interview, it ends its last message with an [END] marker. The runner removes the marker before speaking, finishes the session, and the report evaluation starts.
I will not cover the repository part, because it is not the main point of this challenge. The repositories are all in memory, and you can see the whole code on GitHub (the link is at the end of the article).
The session service is the only business logic of the app. Starting a session just creates it in memory with the state Running. Finishing it is more interesting, because the evaluation can take a few seconds, so I don't make the participant wait for it:
public async Task FinishInterviewSessionAsync(string sessionId)
{
var session = await sessionRepository.GetAsync(sessionId);
if (session is null || session.State != SessionState.Running)
return;
session.State = SessionState.Finished;
session.Duration = timeProvider.GetUtcNow() - session.StartedAt;
await sessionRepository.UpdateAsync(session);
// The evaluation can take a while, so it runs in the background.
// The UI polls GetSessionAsync until the state is Reported.
_ = Task.Run(() => EvaluateAsync(session));
}
The result page shows "Evaluating..." and checks the session every 2 seconds. When the state becomes Reported, the report appears.
This part could be better, but we keep it this way for simplicity. In a real application, the evaluation would go to a queued background job system, so it survives an app restart and can be retried if it fails, and the result would be pushed to the page with SignalR instead of polling.
The interviewer runs on OpenAI through Microsoft.Extensions.AI (IChatClient). Its prompt has two parts:
InterviewInstruction.
The conversation so far is the transcript: what the agent said becomes an assistant message, and what the participant said becomes a user message. Then the reply is streamed back:
public async IAsyncEnumerable<string> RespondAsync(
SkillConfig config, InterviewSession session, [EnumeratorCancellation] CancellationToken ct = default)
{
List<ChatMessage> messages = [new(ChatRole.System, InterviewPromptBuilder.Build(config, session))];
if (session.Transcriptions.Count == 0)
messages.Add(new ChatMessage(ChatRole.User, "(The participant has joined. Start the interview.)"));
foreach (var entry in session.Transcriptions)
{
var role = entry.Speaker == Speaker.Agent ? ChatRole.Assistant : ChatRole.User;
messages.Add(new ChatMessage(role, entry.Text));
}
// The time note goes last, so the model does not miss it.
var timeNote = InterviewPromptBuilder.BuildTimeInstruction(config, session, timeProvider.GetUtcNow());
messages.Add(new ChatMessage(ChatRole.System, timeNote));
var chatClient = chatClients.Get(config.Effort);
await foreach (var update in chatClient.GetStreamingResponseAsync(messages, cancellationToken: ct))
{
if (!string.IsNullOrEmpty(update.Text))
yield return update.Text;
}
}
The time note deserves a comment. While testing, putting the duration and the elapsed time in the system prompt and letting the model do the math didn't work: the agent closed a 20 minute interview in 7 minutes, and at the end it ignored even "Time is up". So the server decides the phase, and the model only gets one short sentence for the current moment:
| Phase | What the agent is told |
|---|---|
| In progress | "5 of 20 minutes have passed (15 minutes left). Do not mention extra time." |
| Extra time starts | "Start your reply by telling the participant that the planned time is over and they have 3 extra minutes." |
| Extra time | "The participant already knows. Ask at most one more question, then close the interview." |
| Time is up | "Time is up. Thank the participant, say goodbye and close the interview now." |
Sending it as the last message, right after the participant's answer, was what made the difference. Inside a long system prompt, the model simply ignored it.
Effort. Each skill chooses how capable the model should be ("Effort": "Low" | "Medium" | "High"), and appsettings.json maps each level to a model:
"OpenAI": {
"Models": { "Low": "gpt-4o-mini", "Medium": "gpt-4.1-mini", "High": "gpt-4.1" }
}
The effort is set per skill, according to what the skill needs: e.g. a light, casual quiz can use a cheaper and faster model, while a deep technical interview gets a stronger one. chatClients.Get(config.Effort) picks it.
The report uses the same model, but instead of streaming text, it asks for structured output. GetResponseAsync<T> generates a JSON schema from the type, OpenAI must answer following that schema, and the result is deserialized back into the type:
var chatClient = chatClients.Get(config.Effort);
var response = await chatClient.GetResponseAsync<ReportResponse>(messages, cancellationToken: ct);
var report = response.Result;
return ReportPromptBuilder.ToResult(config, report.Score, report.Summary, report.StrongPoints, report.ImprovementAreas);
private sealed record ReportResponse(int Score, string Summary, List<string> StrongPoints, List<string> ImprovementAreas);
I don't trust the model with everything: ToResult keeps the score between 1 and MaxPoints, and the label always comes from the skill config ( PointInstructionMap[score].Label), never from the model.
AzureSpeechConverter implements ISpeechConverter with the Azure Speech SDK.
Speech to text. The SDK doesn't work with IAsyncEnumerable, it works with events: Recognizing while the participant speaks (partial text) and Recognized after a silence (final text). A Channel bridges the events to the stream the engine expects:
var segments = Channel.CreateUnbounded<TranscriptionSegment>();
recognizer.Recognizing += (_, e) =>
segments.Writer.TryWrite(new TranscriptionSegment(e.Result.Text, IsFinal: false));
recognizer.Recognized += (_, e) =>
{
if (e.Result.Reason == ResultReason.RecognizedSpeech && !string.IsNullOrWhiteSpace(e.Result.Text))
segments.Writer.TryWrite(new TranscriptionSegment(e.Result.Text, IsFinal: true));
};
await recognizer.StartContinuousRecognitionAsync();
await foreach (var segment in segments.Reader.ReadAllAsync(ct))
yield return segment;
The microphone audio is written into the recognizer's push stream in the background while the segments are read. The silence that ends an answer is one setting, SilenceTimeoutMs (2000 ms in appsettings.json). With 1 second the agent cut me while I was thinking, and 2 seconds felt natural.
Text to speech. Creating a new SpeechSynthesizer for every sentence opens a new connection to Azure each time: about 1.2 seconds per sentence in my tests. So there is one synthesizer per voice, created once and reused:
private SpeechSynthesizer GetSynthesizer(string voiceName) =>
synthesizers.GetOrAdd(voiceName, CreateSynthesizer);
After the first sentence, each one takes about 0.2 seconds. That second saved in every reply is the difference between a conversation and a walkie-talkie.
The runner is the heart of the engine. It listens to the transcription all the time and reacts to each segment:
// The agent opens the interview.
StartAgentTurn(state);
var participantAudio = channel.ReadParticipantAudioAsync(state.InterviewToken);
var segments = speech.ToTranscriptionAsync(participantAudio, state.InterviewToken);
await foreach (var segment in segments.WithCancellation(state.InterviewToken))
{
if (segment.IsFinal)
await OnParticipantFinishedAsync(state, segment.Text, onPartialTranscript);
else
await OnParticipantSpeakingAsync(state, segment.Text, onPartialTranscript);
}
While the participant speaks, the caption is updated and the agent stops talking. That's the interruption:
private async Task OnParticipantSpeakingAsync(
RealtimeInterviewState state, string partialText, Func<string, Task>? onPartialTranscript)
{
if (onPartialTranscript is not null)
await onPartialTranscript(partialText);
if (state.AgentInterrupted)
return;
state.AgentInterrupted = true;
await state.AgentTurnCts!.CancelAsync();
await state.Channel.StopAgentAudioAsync();
}
Cancelling the agent's turn stops everything at once: the OpenAI stream, the speech synthesis and the audio still queued in the browser. This is why the CancellationToken matters here, and not in the repositories.
When the participant finishes, the phrase goes to the transcript and a new agent turn starts in the background. The turn streams the agent's reply, a small SentenceSplitter cuts it into sentences, and each sentence is spoken as soon as it is complete, while the model is still writing the next one. The participant hears the first sentence after about a second.
One detail: the runner is a singleton, but each interview has its own state (the current agent turn, its cancellation, whether it was interrupted). That state lives in a small class, RealtimeInterviewState, created at the beginning of RunAsync.
The voice has its own WebSocket, separate from the SignalR connection Blazor uses for the UI. Blazor's JS interop could carry the audio too, but a plain WebSocket is much easier to follow (and honestly, JS interop always felt a bit weird to me). The browser just sends and receives frames, with a tiny protocol:
| Direction | Frame | Content |
|---|---|---|
| browser → server | binary | Microphone audio, 16-bit PCM at 16 kHz, 100 ms per frame |
| server → browser | binary | Agent voice, 16-bit PCM at 24 kHz, one sentence per frame |
| server → browser | text | {"type":"caption","text":"..."} |
| server → browser | text | {"type":"stop"} : the participant interrupted |
| server → browser | text | {"type":"ended"} : the agent closed the interview |
| server → browser | text | {"type":"error","message":"..."} |
On the server, it's a normal ASP.NET Core endpoint. It accepts the socket, wraps it in a WebSocketAudioChannel (our IAudioChannel) and runs the interview:
app.UseWebSockets();
app.Map("/interview/{sessionId}/voice", HandleAsync);
// inside HandleAsync
using var socket = await context.WebSockets.AcceptWebSocketAsync();
var channel = new WebSocketAudioChannel(socket);
await runner.RunAsync(sessionId, config, channel, channel.SendCaptionAsync, interviewCts.Token);
WebSocketAudioChannel turns binary frames into AudioChunk s and events into small JSON messages. One trap: a WebSocket allows only one send at a time, and the agent's voice and the events are sent from different tasks, so the sends go through a SemaphoreSlim.
In the browser, the voice is a web component: a custom HTML tag with its own JavaScript. The Blazor page only renders the tag:
<voice-panel session-id="@SessionId" agent-name="@config.AgentName"></voice-panel>
and voice-panel.js does the rest: the "Start talking" button, the microphone, the socket and the speaker.
this.socket = new WebSocket(`${protocol}//${location.host}/interview/${sessionId}/voice`);
this.socket.binaryType = "arraybuffer";
this.socket.onopen = () => this.startMicrophone();
this.socket.onmessage = event => this.onMessage(event.data);
The microphone goes through an AudioWorklet that converts the audio to 16 kHz PCM and sends a frame every 100 ms. The agent's audio is played gapless, each sentence scheduled right after the previous one. When you leave the page, the tag is removed, the socket closes, and the server stops the interview loop. No JS interop at all.
A nice bonus: in the browser DevTools (Network → WS → voice) you can watch every frame of the conversation.
If you have built something similar, I would love to hear about your experience.
One thing worth trying is keeping the speech-to-text and text-to-speech on the client side, in JavaScript. I tried it, but it didn't work well, so I decided to move it to the backend.
The whole code is on GitHub: github.com/hugoj0s3/SkillsValidator
OpenAI
Azure Speech
eastus or westeurope, not the display name. The key only works with its own region. The Clone the repository:
git clone https://github.com/hugoj0s3/SkillsValidator
cd SkillsValidator
The keys go in user-secrets, so they never end up in the code or in git:
dotnet user-secrets set "OpenAI:ApiKey" "<your OpenAI key>" --project src/SkillsValidator.Web
dotnet user-secrets set "AzureSpeech:Key" "<your Azure Speech key>" --project src/SkillsValidator.Web
dotnet user-secrets set "AzureSpeech:Region" "<your region, e.g. eastus>" --project src/SkillsValidator.Web
Start the app:
dotnet run --project src/SkillsValidator.Web
Open http://localhost:5101, choose a skill, type your name and click Start interview. On the interview page, click Start talking and allow the microphone.
If a key or a skill file is missing or wrong, the start page tells you exactly what to fix.
src/SkillsValidator.Web/skills and restart the app.OpenAI:Models in appsettings.json maps each effort level (Low, Medium, High) to a model.AzureSpeech:SilenceTimeoutMs is how long the participant can be silent before the agent answers (2000 ms by default).AzureSpeech:Voices maps each gender and voice type to an Azure voice.