{"slug": "building-a-skill-interview-with-ai", "title": "Building a Skill Interview with AI", "summary": "A developer built a voice-based AI skill assessment app in which an AI agent interviews a participant about a chosen skill and produces a scored report card. Skills are defined entirely in JSON files (including agent name, tone, voice, model effort, duration and per-point scoring requirements), so new skills can be added without code changes, and the OpenAI and Azure Speech components sit behind interfaces so providers can be swapped. The developer reports that calibrating the prompt, particularly ending the interview at the right time, was the hardest part.", "body_md": "Here's the idea: build an app that checks if someone knows their stuff using an interview. You pick a skill, type your name, and off you go. An AI agent chats with you by voice, asks questions, and at the end you get a report card. Hopefully with more compliments than trauma.\n\nI've seen some applications that evaluate interview participants with AI, so I decided to build one out of curiosity. It was an interesting journey. The trickiest part was calibrating the prompt, especially to end the interview at the right time.\n\nLet's get the requirements first, so we have a clear picture of what to implement, before I start the architecture, the code, drawing boxes and arrows, and pretending I know what I'm doing lol.\n\nEvery skill is a JSON file. Adding a new skill means adding a file: no code changes. A skill defines:\n\n| Field | Purpose | \n|---|---|\n| `Title` | Name shown to the participant, e.g. \"SQL\" | \n| `InterviewInstruction` | What the agent should ask and which topics matter most | \n| `AgentName` ,`AgentTone` | Who the agent is and how it behaves (friendly, professional, like a Jedi master...) | \n| `AgentVoiceGender` ,`AgentVoiceType` | The agent's voice (Male/Female; HighPitched, Neutral, Deep, Warm) | \n| `Effort` | How capable the AI model should be: Low, Medium or High | \n| `ReportInstruction` | How to evaluate and what to put in the report | \n| `MaxPoints` | The highest possible score | \n| `InterviewDurationInMinutes` ,`ExtraInterviewDurationInMinutes` | Planned duration and extra time | \n| `PointInstructionMap` | A label and the minimum requirements for **every** point from 1 to`MaxPoints` | \n\nFor example, a Star Wars skill can use fun level names:\n\n```\n{\n  \"Id\": \"star-wars\",\n  \"Title\": \"Star Wars Lore\",\n  \"InterviewInstruction\": \"Test the participant's knowledge of the Star Wars universe...\",\n  \"AgentTone\": \"Warm and playful, like a wise old Jedi master.\",\n  \"AgentName\": \"Master Oren\",\n  \"AgentVoiceGender\": \"Male\",\n  \"AgentVoiceType\": \"Deep\",\n  \"Effort\": \"Low\",\n  \"ReportInstruction\": \"Evaluate breadth and depth of lore knowledge.\",\n  \"MaxPoints\": 4,\n  \"InterviewDurationInMinutes\": 10,\n  \"ExtraInterviewDurationInMinutes\": 2,\n  \"PointInstructionMap\": {\n    \"1\": { \"Label\": \"Youngling\", \"Requirements\": \"Knows the main characters.\" },\n    \"2\": { \"Label\": \"Padawan\", \"Requirements\": \"Knows the plot of the main films.\" },\n    \"3\": { \"Label\": \"Jedi Knight\", \"Requirements\": \"Knows the history of the Jedi and the Sith.\" },\n    \"4\": { \"Label\": \"Jedi Master\", \"Requirements\": \"Knows deep lore, including series and books.\" }\n  }\n}\n```\n\nIf a skill file is invalid (for example, a missing point in `PointInstructionMap`), the app starts with a clear message saying what to fix.\n\nThis is a sample for an article, so some things are deliberately left out:\n\nThe interview logic is agnostic and doesn't depend on a specific AI vendor.\n\nThe interviewer and the report agent (**OpenAI**) and the speech-to-text/text-to-speech service (** Azure Speech**) are each behind an interface, in their own project, so that they can be replaced by different providers.\n\nLet's build the abstractions first: the backend components, and how the interview engine handles everything the participant says so the conversation flows as naturally as possible.\n\nThe engine is split into two parts:\n\nThe participant's voice is transcribed and sent to the agent; then the agent writes its reply, and the reply is synthesized into speech with the voice from the config.\n\nEvery box in the bottom row is an interface, implemented in its own project (`SkillsValidator.Engine.OpenAI`, `SkillsValidator.Engine.AzureSpeech`). The engine never references a vendor SDK.\n\nSkills are read-only: they come from the JSON files.\n\n```\npublic interface ISkillConfigRepository\n{\n    Task<IReadOnlyList<SkillConfig>> GetAllAsync();\n    Task<SkillConfig?> GetAsync(string configId);\n    SkillConfigLoadResult GetLoadResult(); // startup validation: the errors to show if a file is invalid\n}\n```\n\nAn interview is a **session**. It goes through three states: `Running` → `Finished` (evaluation pending) → `Reported` (the report is ready).\n\n```\npublic interface ISessionInterviewService\n{\n    Task<string> StartInterviewSessionAsync(string participantName, SkillConfig config);\n    Task FinishInterviewSessionAsync(string sessionId); // also starts the evaluation\n    Task<InterviewSession?> GetSessionAsync(string sessionId);\n}\n```\n\n`GetSessionAsync` returns the whole picture: the participant, the state, the transcript and, once `Reported`, the result. Behind it, three small repositories (session, transcript, result) keep everything in memory. The transcript is **append-only**: an interview only ever adds new lines.\n\nWhen the interview finishes, a second agent evaluates it:\n\n```\npublic interface IReportEvaluator\n{\n    Task<SessionReportResult> EvaluateAsync(\n        SkillConfig config, IReadOnlyList<TranscriptEntry> transcript, CancellationToken ct = default);\n}\n```\n\nFour interfaces make the voice conversation work.\n\n**1. Hearing and speaking:** `ISpeechConverter` turns audio into text and text into audio.\n\n```\npublic interface ISpeechConverter\n{\n    // Microphone audio in, text out: partial text while the participant speaks, final text after a short silence.\n    IAsyncEnumerable<TranscriptionSegment> ToTranscriptionAsync(\n        IAsyncEnumerable<AudioChunk> speech, CancellationToken ct = default);\n\n    // Text in, the agent's voice out.\n    IAsyncEnumerable<AudioChunk> ToSpeechAsync(\n        string text, VoiceGender gender, VoiceType type, string? speakingStyle, CancellationToken ct = default);\n}\n\npublic sealed record AudioChunk(byte[] Data, string Format, int SampleRate);\npublic sealed record TranscriptionSegment(string Text, bool IsFinal);\n```\n\nNotice that everything is an `IAsyncEnumerable`: audio and text **flow** through the system in small pieces. The engine never waits for a whole answer before doing the next step.\n\n**2. Thinking:** `IRealtimeInterviewAgent` is the interviewer. Given the skill and the conversation so far, it streams its next reply as text fragments.\n\n```\npublic interface IRealtimeInterviewAgent\n{\n    IAsyncEnumerable<string> RespondAsync(SkillConfig config, InterviewSession session, CancellationToken ct = default);\n}\n```\n\n**3. The pipe to the browser:** `IAudioChannel` receives the participant's microphone audio and plays the agent's voice.\n\n```\npublic interface IAudioChannel\n{\n    IAsyncEnumerable<AudioChunk> ReadParticipantAudioAsync(CancellationToken ct);\n    Task PlayAgentAudioAsync(IAsyncEnumerable<AudioChunk> audio, CancellationToken ct);\n    Task StopAgentAudioAsync(); // the participant interrupted\n}\n```\n\nThe engine doesn't care how the audio travels. In our app it's a WebSocket.\n\n**4. The conductor:** `IRealtimeInterviewRunner` connects the other three and runs the conversation until it ends.\n\n```\npublic interface IRealtimeInterviewRunner\n{\n    Task RunAsync(string sessionId, SkillConfig config, IAudioChannel channel,\n        Func<string, Task>? onPartialTranscript, // live caption\n        CancellationToken ct);\n}\n```\n\nThe runner listens all the time. The speech converter sends it two kinds of events, and each one gets a different reaction:\n\n**While the participant is speaking** (a *partial* segment):\n\n`StopAgentAudioAsync` silences the speaker. This is what makes **When the participant has finished** (a *final* segment, after about 2 seconds of silence):\n\nSpeaking sentence by sentence is the key to a natural flow. The participant hears the first sentence after roughly a second, instead of waiting for the whole reply to be written and then spoken.\n\n**Keeping the agent on track.** Two details make the conversation feel less robotic:\n\n**Ending.** When the agent decides to close the interview, it ends its last message with an `[END]` marker. The runner removes the marker before speaking, finishes the session, and the report evaluation starts.\n\nI will not cover the repository part, because it is not the main point of this challenge. The repositories are all in memory, and you can see the whole code on GitHub (the link is at the end of the article).\n\nThe session service is the only business logic of the app. Starting a session just creates it in memory with the state `Running`. Finishing it is more interesting, because the evaluation can take a few seconds, so I don't make the participant wait for it:\n\n``` js\npublic async Task FinishInterviewSessionAsync(string sessionId)\n{\n    var session = await sessionRepository.GetAsync(sessionId);\n    if (session is null || session.State != SessionState.Running)\n        return;\n\n    session.State = SessionState.Finished;\n    session.Duration = timeProvider.GetUtcNow() - session.StartedAt;\n    await sessionRepository.UpdateAsync(session);\n\n    // The evaluation can take a while, so it runs in the background.\n    // The UI polls GetSessionAsync until the state is Reported.\n    _ = Task.Run(() => EvaluateAsync(session));\n}\n```\n\nThe result page shows \"Evaluating...\" and checks the session every 2 seconds. When the state becomes `Reported`, the report appears.\n\nThis part could be better, but we keep it this way for simplicity. In a real application, the evaluation would go to a queued background job system, so it survives an app restart and can be retried if it fails, and the result would be pushed to the page with SignalR instead of polling.\n\nThe interviewer runs on OpenAI through **Microsoft.Extensions.AI** (`IChatClient`). Its prompt has two parts:\n\n`InterviewInstruction`.\nThe conversation so far is the transcript: what the agent said becomes an `assistant` message, and what the participant said becomes a `user` message. Then the reply is streamed back:\n\n```\npublic async IAsyncEnumerable<string> RespondAsync(\n    SkillConfig config, InterviewSession session, [EnumeratorCancellation] CancellationToken ct = default)\n{\n    List<ChatMessage> messages = [new(ChatRole.System, InterviewPromptBuilder.Build(config, session))];\n\n    if (session.Transcriptions.Count == 0)\n        messages.Add(new ChatMessage(ChatRole.User, \"(The participant has joined. Start the interview.)\"));\n\n    foreach (var entry in session.Transcriptions)\n    {\n        var role = entry.Speaker == Speaker.Agent ? ChatRole.Assistant : ChatRole.User;\n        messages.Add(new ChatMessage(role, entry.Text));\n    }\n\n    // The time note goes last, so the model does not miss it.\n    var timeNote = InterviewPromptBuilder.BuildTimeInstruction(config, session, timeProvider.GetUtcNow());\n    messages.Add(new ChatMessage(ChatRole.System, timeNote));\n\n    var chatClient = chatClients.Get(config.Effort);\n    await foreach (var update in chatClient.GetStreamingResponseAsync(messages, cancellationToken: ct))\n    {\n        if (!string.IsNullOrEmpty(update.Text))\n            yield return update.Text;\n    }\n}\n```\n\nThe time note deserves a comment. While testing, putting the duration and the elapsed time in the system prompt and letting the model do the math didn't work: the agent closed a 20 minute interview in 7 minutes, and at the end it ignored even \"Time is up\". So the **server** decides the phase, and the model only gets one short sentence for the current moment:\n\n| Phase | What the agent is told | \n|---|---|\n| In progress | \"5 of 20 minutes have passed (15 minutes left). Do not mention extra time.\" | \n| Extra time starts | \"Start your reply by telling the participant that the planned time is over and they have 3 extra minutes.\" | \n| Extra time | \"The participant already knows. Ask at most one more question, then close the interview.\" | \n| Time is up | \"Time is up. Thank the participant, say goodbye and close the interview now.\" | \n\nSending it as the **last message**, right after the participant's answer, was what made the difference. Inside a long system prompt, the model simply ignored it.\n\n**Effort.** Each skill chooses how capable the model should be (`\"Effort\": \"Low\" | \"Medium\" | \"High\"`), and `appsettings.json` maps each level to a model:\n\n```\n\"OpenAI\": {\n  \"Models\": { \"Low\": \"gpt-4o-mini\", \"Medium\": \"gpt-4.1-mini\", \"High\": \"gpt-4.1\" }\n}\n```\n\nThe effort is set per skill, according to what the skill needs: e.g. a light, casual quiz can use a cheaper and faster model, while a deep technical interview gets a stronger one. `chatClients.Get(config.Effort)` picks it.\n\nThe report uses the same model, but instead of streaming text, it asks for **structured output**. `GetResponseAsync<T>` generates a JSON schema from the type, OpenAI must answer following that schema, and the result is deserialized back into the type:\n\n``` js\nvar chatClient = chatClients.Get(config.Effort);\nvar response = await chatClient.GetResponseAsync<ReportResponse>(messages, cancellationToken: ct);\n\nvar report = response.Result;\nreturn ReportPromptBuilder.ToResult(config, report.Score, report.Summary, report.StrongPoints, report.ImprovementAreas);\n\nprivate sealed record ReportResponse(int Score, string Summary, List<string> StrongPoints, List<string> ImprovementAreas);\n```\n\nI don't trust the model with everything: `ToResult` keeps the score between 1 and `MaxPoints`, and the label always comes from the skill config (` PointInstructionMap[score].Label`), never from the model.\n\n`AzureSpeechConverter` implements `ISpeechConverter` with the Azure Speech SDK.\n\n**Speech to text.** The SDK doesn't work with `IAsyncEnumerable`, it works with events: `Recognizing` while the participant speaks (partial text) and `Recognized` after a silence (final text). A `Channel` bridges the events to the stream the engine expects:\n\n``` js\nvar segments = Channel.CreateUnbounded<TranscriptionSegment>();\n\nrecognizer.Recognizing += (_, e) =>\n    segments.Writer.TryWrite(new TranscriptionSegment(e.Result.Text, IsFinal: false));\n\nrecognizer.Recognized += (_, e) =>\n{\n    if (e.Result.Reason == ResultReason.RecognizedSpeech && !string.IsNullOrWhiteSpace(e.Result.Text))\n        segments.Writer.TryWrite(new TranscriptionSegment(e.Result.Text, IsFinal: true));\n};\n\nawait recognizer.StartContinuousRecognitionAsync();\n\nawait foreach (var segment in segments.Reader.ReadAllAsync(ct))\n    yield return segment;\n```\n\nThe microphone audio is written into the recognizer's push stream in the background while the segments are read. The silence that ends an answer is one setting, `SilenceTimeoutMs` (2000 ms in `appsettings.json`). With 1 second the agent cut me while I was thinking, and 2 seconds felt natural.\n\n**Text to speech.** Creating a new `SpeechSynthesizer` for every sentence opens a new connection to Azure each time: about **1.2 seconds per sentence** in my tests. So there is one synthesizer per voice, created once and reused:\n\n``` js\nprivate SpeechSynthesizer GetSynthesizer(string voiceName) =>\n    synthesizers.GetOrAdd(voiceName, CreateSynthesizer);\n```\n\nAfter the first sentence, each one takes about **0.2 seconds**. That second saved in every reply is the difference between a conversation and a walkie-talkie.\n\nThe runner is the heart of the engine. It listens to the transcription all the time and reacts to each segment:\n\n``` js\n// The agent opens the interview.\nStartAgentTurn(state);\n\nvar participantAudio = channel.ReadParticipantAudioAsync(state.InterviewToken);\nvar segments = speech.ToTranscriptionAsync(participantAudio, state.InterviewToken);\n\nawait foreach (var segment in segments.WithCancellation(state.InterviewToken))\n{\n    if (segment.IsFinal)\n        await OnParticipantFinishedAsync(state, segment.Text, onPartialTranscript);\n    else\n        await OnParticipantSpeakingAsync(state, segment.Text, onPartialTranscript);\n}\n```\n\nWhile the participant speaks, the caption is updated and the agent stops talking. That's the interruption:\n\n```\nprivate async Task OnParticipantSpeakingAsync(\n    RealtimeInterviewState state, string partialText, Func<string, Task>? onPartialTranscript)\n{\n    if (onPartialTranscript is not null)\n        await onPartialTranscript(partialText);\n\n    if (state.AgentInterrupted)\n        return;\n\n    state.AgentInterrupted = true;\n    await state.AgentTurnCts!.CancelAsync();\n    await state.Channel.StopAgentAudioAsync();\n}\n```\n\nCancelling the agent's turn stops everything at once: the OpenAI stream, the speech synthesis and the audio still queued in the browser. This is why the `CancellationToken` matters here, and not in the repositories.\n\nWhen the participant finishes, the phrase goes to the transcript and a new agent turn starts in the background. The turn streams the agent's reply, a small `SentenceSplitter` cuts it into sentences, and **each sentence is spoken as soon as it is complete**, while the model is still writing the next one. The participant hears the first sentence after about a second.\n\nOne detail: the runner is a singleton, but each interview has its own state (the current agent turn, its cancellation, whether it was interrupted). That state lives in a small class, `RealtimeInterviewState`, created at the beginning of `RunAsync`.\n\nThe voice has its own WebSocket, separate from the SignalR connection Blazor uses for the UI. Blazor's JS interop could carry the audio too, but a plain WebSocket is much easier to follow (and honestly, JS interop always felt a bit weird to me). The browser just sends and receives frames, with a tiny protocol:\n\n| Direction | Frame | Content | \n|---|---|---|\n| browser → server | binary | Microphone audio, 16-bit PCM at 16 kHz, 100 ms per frame | \n| server → browser | binary | Agent voice, 16-bit PCM at 24 kHz, one sentence per frame | \n| server → browser | text | `{\"type\":\"caption\",\"text\":\"...\"}` | \n| server → browser | text | `{\"type\":\"stop\"}` : the participant interrupted | \n| server → browser | text | `{\"type\":\"ended\"}` : the agent closed the interview | \n| server → browser | text | `{\"type\":\"error\",\"message\":\"...\"}` | \n\nOn the server, it's a normal ASP.NET Core endpoint. It accepts the socket, wraps it in a `WebSocketAudioChannel` (our `IAudioChannel`) and runs the interview:\n\n```\napp.UseWebSockets();\napp.Map(\"/interview/{sessionId}/voice\", HandleAsync);\n\n// inside HandleAsync\nusing var socket = await context.WebSockets.AcceptWebSocketAsync();\nvar channel = new WebSocketAudioChannel(socket);\n\nawait runner.RunAsync(sessionId, config, channel, channel.SendCaptionAsync, interviewCts.Token);\n```\n\n`WebSocketAudioChannel` turns binary frames into `AudioChunk` s and events into small JSON messages. One trap: a WebSocket allows only one send at a time, and the agent's voice and the events are sent from different tasks, so the sends go through a `SemaphoreSlim`.\n\nIn the browser, the voice is a **web component**: a custom HTML tag with its own JavaScript. The Blazor page only renders the tag:\n\n```\n<voice-panel session-id=\"@SessionId\" agent-name=\"@config.AgentName\"></voice-panel>\n```\n\nand `voice-panel.js` does the rest: the \"Start talking\" button, the microphone, the socket and the speaker.\n\n```\nthis.socket = new WebSocket(`${protocol}//${location.host}/interview/${sessionId}/voice`);\nthis.socket.binaryType = \"arraybuffer\";\nthis.socket.onopen = () => this.startMicrophone();\nthis.socket.onmessage = event => this.onMessage(event.data);\n```\n\nThe microphone goes through an `AudioWorklet` that converts the audio to 16 kHz PCM and sends a frame every 100 ms. The agent's audio is played gapless, each sentence scheduled right after the previous one. When you leave the page, the tag is removed, the socket closes, and the server stops the interview loop. No JS interop at all.\n\nA nice bonus: in the browser DevTools (Network → WS → `voice`) you can watch every frame of the conversation.\n\nIf you have built something similar, I would love to hear about your experience.\n\nOne thing worth trying is keeping the speech-to-text and text-to-speech on the client side, in JavaScript. I tried it, but it didn't work well, so I decided to move it to the backend.\n\nThe whole code is on GitHub: [github.com/hugoj0s3/SkillsValidator](https://github.com/hugoj0s3/SkillsValidator)\n\n**OpenAI**\n\n**Azure Speech**\n\n`eastus` or `westeurope`, not the display name. The key only works with its own region. The Clone the repository:\n\n```\ngit clone https://github.com/hugoj0s3/SkillsValidator\ncd SkillsValidator\n```\n\nThe keys go in **user-secrets**, so they never end up in the code or in git:\n\n```\ndotnet user-secrets set \"OpenAI:ApiKey\" \"<your OpenAI key>\" --project src/SkillsValidator.Web\ndotnet user-secrets set \"AzureSpeech:Key\" \"<your Azure Speech key>\" --project src/SkillsValidator.Web\ndotnet user-secrets set \"AzureSpeech:Region\" \"<your region, e.g. eastus>\" --project src/SkillsValidator.Web\n```\n\nStart the app:\n\n```\ndotnet run --project src/SkillsValidator.Web\n```\n\nOpen [http://localhost:5101](http://localhost:5101), choose a skill, type your name and click **Start interview**. On the interview page, click **Start talking** and allow the microphone.\n\nIf a key or a skill file is missing or wrong, the start page tells you exactly what to fix.\n\n`src/SkillsValidator.Web/skills` and restart the app.`OpenAI:Models` in `appsettings.json` maps each effort level (Low, Medium, High) to a model.`AzureSpeech:SilenceTimeoutMs` is how long the participant can be silent before the agent answers (2000 ms by default).`AzureSpeech:Voices` maps each gender and voice type to an Azure voice.", "url": "https://wpnews.pro/news/building-a-skill-interview-with-ai", "canonical_source": "https://dev.to/hugo_jose_9/building-a-skill-interview-with-ai-4mjl", "published_at": "2026-09-28 08:21:54+00:00", "updated_at": "2026-09-28 08:49:34.531368+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "ai-products", "natural-language-processing", "developer-tools"], "entities": ["OpenAI", "Azure Speech", "SkillsValidator.Engine.OpenAI", "SkillsValidator.Engine.AzureSpeech"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-a-skill-interview-with-ai", "markdown": "https://wpnews.pro/news/building-a-skill-interview-with-ai.md", "text": "https://wpnews.pro/news/building-a-skill-interview-with-ai.txt", "jsonld": "https://wpnews.pro/news/building-a-skill-interview-with-ai.jsonld"}}