HomeCase studiesA Conversational AI Tutoring Platform for Professional...

Case study · Seattle, United States

A Conversational AI Tutoring Platform for Professional Education

Professional education technology — structured, case-study-based training for working practitioners

The most valuable moment in professional training is when an expert reads a learner's reasoning and tells them exactly where it goes wrong. It is also the least scalable. Our team built a learning platform that reproduces that interaction at scale: learners work through case-study modules, answer structured questions, and then converse with an AI tutor whose responses are grounded in that specific module's teaching material and that specific question's guidance. The architectural decision that makes it work is that the prompts are content, not code — subject-matter experts edit the tutor's persona, the module context and the per-question guidance directly through the admin interface, and every conversation turn is stored with the exact prompt that produced it.

Industry
Professional education technology — structured, case-study-based training for working practitioners
Solution
A web-based learning platform with an integrated conversational AI tutor, a full curriculum and prompt authoring back office, and a supporting REST API.
Platforms
Web & API
Stack
Laravel · PHP · Blade · jQuery · Bootstrap · Tailwind
Location
Seattle, United States · delivered remotely
Role
Engineering, with the team behind for mobile and QA

Delivered remotely for a business based in Seattle, United States. Names withheld by agreement.

Project Overview

Industry: Professional education technology — structured, case-study-based training for working practitioners.

Type of solution: A web-based learning platform with an integrated conversational AI tutor, a full curriculum and prompt authoring back office, and a supporting REST API.

Business context: Advanced professional training depends on expert feedback, and expert feedback does not scale. A senior practitioner reviewing a learner's reasoning one-to-one is the highest-value part of any professional course and the part that caps enrolment. Conventional e-learning solves this by removing the interaction: automated marking can tell a learner they are wrong but cannot tell them why, cannot probe the reasoning behind an answer, and cannot adapt when the misunderstanding is different from the one the question anticipated.

General users: Learners working through course modules, subject-matter experts authoring curriculum and tutoring prompts, and platform administrators.

General purpose: To deliver structured, guided professional training in which every learner receives a substantive, context-grounded conversation about their own reasoning — and to put control of that conversation's quality in the hands of the educators rather than the engineers.

The Business Challenge

Expert feedback is the product, and it does not scale. The instructional value sits in a 1:1 conversation with someone who knows the subject. Those people are scarce, expensive, and busy. Any platform that cannot reproduce that conversation is selling a lesser product.

Automated marking cannot diagnose reasoning. A multiple-choice engine knows whether an answer matches a key. It does not know that the learner reached the right answer for the wrong reason, or the wrong answer through sound logic with one faulty assumption — which is the feedback that actually changes how someone thinks.

A general-purpose model is not a subject-matter tutor. An unguided language model will answer from general knowledge, in a generic voice, at whatever length it chooses. For a professional course, responses have to be grounded in this module's material, aligned with this question's teaching intent, delivered in a consistent voice, and bounded in length and scope.

The people who can fix the tutor are not engineers. When a tutor response is unhelpful, the correction is almost always a prompt change — more context, a sharper guardrail, a clearer instruction. If every prompt change requires a developer and a deployment, iteration slows to the speed of the release cycle and the educators lose ownership of the thing they are best placed to improve.

AI behaviour has to be explainable after the fact. When a learner reports a poor tutor response, the team needs to know exactly what was sent to the model at that moment. If prompts have since been edited, and only the response was stored, that question is unanswerable.

Answers come in many shapes but the conversation has one. Text responses, single and grouped selections, image-based matching exercises and rating scales all have to enter the same conversational pipeline and read as something a person actually said.

Every model call costs money and time. Unbounded follow-up chat is unbounded expense, and a third-party API call inside a page interaction is latency the learner feels.

Learning is not one sitting. Learners leave mid-module and come back. Progress, answers and the full conversation have to be exactly where they left them.

Our Approach

Make the prompts editable content, owned by educators. The tutor's persona and guardrails live on the module. The module's teaching material lives on the module. Each question carries its own guidance and its own model-facing phrasing. All of it is stored in the database and edited through purpose-built prompt editors in the admin interface, separate from the ordinary content forms and behind an additional access gate. Improving the tutor is a content change, not a release.

Compose the prompt in layers with clear ownership. Persona and behavioural guardrails form the system layer. The module context, the question's guidance, the learner's answer and the prior conversation form a structured, labelled user layer. Each layer has one owner and one editing surface, which keeps prompt engineering tractable as the curriculum grows.

Store the prompt with the response, every turn. Each conversation turn persists its role, its content, the exact system prompt and user prompt used to generate it, and the raw provider payload. When a tutor response needs to be explained, the team reads exactly what was sent — not an approximation reconstructed from prompts that have since changed.

Treat rich text as a prompt engineering problem. Teaching material is authored in a WYSIWYG editor and stored as markup. Sending that markup to a model wastes tokens and degrades output; stripping tags naively destroys the structure the meaning lives in. We built a dedicated normalisation routine that parses the content properly, decodes entities, normalises whitespace, extracts the visible text, and then reconstructs paragraph, section and list structure so the model receives clean, well-formed prose with its organisation intact.

Normalise every answer type into natural language before the model sees it. Selections, groupings, ratings and free text are all rendered into a sentence describing what the learner actually expressed, so the conversation reads coherently no matter which input widget produced it.

Bound the conversation deliberately. A per-learner, per-question follow-up allowance is tracked in the database and reflected in the interface, so the technical limit and the tutor's own conversational limit agree. Where a step requires only acknowledgement and there is nothing to evaluate, the model is not called at all.

Make everything resumable. Completed questions, prior answers, section-level and overall progress, and the full conversation for every question are rehydrated from the database on load, so a learner returning after a week continues rather than restarts.

Give prompt iteration somewhere to run. The delivery pipeline creates environments from branches — including per-feature and experiment environments — so a curriculum or prompt variant can be trialled in a live environment without disturbing the main staging line.

The Solution

Module and curriculum authoring. Modules are composed of sections, each containing an ordered sequence of learner steps, with drag-and-drop reordering, rich-text content, image upload and presentation types controlling how each step is rendered.

Layered prompt management. Each module carries a tutor persona prompt and a body of teaching context. Each question carries its own guidance and its own model-facing phrasing. Both are edited through dedicated prompt editors, with the original engineered prompts retained in configuration as a fallback for modules not yet given their own.

Multiple answer formats. Free text, single selection, grouped selection, image-based grouped matching, rating scales, and an acknowledgement step that completes without evaluation — each with its own authoring form and its own rendering.

Grounded tutoring conversation. When a learner submits an answer, the platform assembles the persona, the module context, the question guidance, the learner's answer rendered as natural language and the prior conversation, and returns a response that corrects misunderstandings, probes reasoning or asks a clarifying question.

Bounded follow-up. Learners can continue the conversation within a question up to a tracked allowance, with the remaining count surfaced in the interface and enforced server-side.

Full conversation audit trail. Every turn is stored with the exact prompts that produced it and the raw provider response, making tutor behaviour reviewable and diagnosable.

Progress and resumption. Section-level and overall progress indicators, a step counter, completion tracking, and full restoration of answers and conversation on return.

Answer reset. A learner can reset a step, clearing the answer, its detail records and its conversation, and restoring the follow-up allowance.

Administrative back office. Role-separated administration covering modules, questions, question and answer types, answer options, learners, administrators, branding and platform settings, with prompt and curriculum management behind an additional gate.

Supporting API. A token-authenticated REST interface covering registration, authentication, password recovery and profile management for a companion client.

Installable web experience. Progressive Web App support so the learner application can be installed and launched like a native application.

Key Features

Database-managed, layered AI prompts Persona, module context and per-question guidance are stored as content and edited through dedicated admin screens, so the people who understand the subject matter control the quality of the tutoring without a code change.

Full prompt-and-response audit trail Every conversation turn persists the exact system prompt, the exact user prompt and the raw provider payload alongside the message itself, so any tutor response can be explained after the fact even after prompts have changed.

Rich-text to prompt-text normalisation A dedicated routine converts WYSIWYG-authored teaching content into clean, structurally faithful plain text for the model — preserving numbered sections, labelled headings and lists rather than flattening them.

Answer-type normalisation into a single conversational pipeline Selections, groupings, image matching, ratings and free text are all rendered into natural language before reaching the model, so one pipeline serves every question format.

Bounded, quota-managed conversation A per-learner, per-question follow-up allowance is enforced server-side and reflected in the interface, keeping cost predictable and conversations purposeful.

Evaluation-free steps that skip the model entirely Acknowledgement steps complete without an API call, because there is nothing to assess — cost control by design rather than by throttling.

Resumable learning state Progress, answers and complete per-question conversations are restored from the database on load, so learners continue exactly where they stopped.

Two-level curriculum structure with ordered steps Sections and their ordered child steps drive both the sidebar navigation and the progress model, with reordering handled from the admin interface.

Separated prompt-engineering access Curriculum and prompt management sits behind an additional gate beyond ordinary administrative access, distinguishing prompt engineering from routine user administration.

Branch-driven environments for curriculum experiments The delivery pipeline provisions environments from branches, including feature and experiment lines, so prompt and curriculum variants can be trialled live without disturbing the main staging environment.

Technical Architecture

Learner application. Server-rendered templates with a sectioned sidebar, step navigation and progress indicators, enhanced with asynchronous requests for answer submission and conversation so the page never reloads mid-exercise. Typing indicators manage perceived latency while the model responds. Full state is rendered from the database on load, making the experience resumable without client-side persistence.

Application layer. A framework monolith with three controller trees — learner, administrative and API — over shared base controllers. Role middleware separates learner and administrative routes; an additional session-check middleware gates the curriculum and prompt management area.

Conversation pipeline. The core flow is deliberately linear and lives in one place: normalise the submitted answer into natural language, load the module persona and context and the question guidance, convert all rich-text content to clean prompt text, compose the layered user prompt with prior history, persist the learner turn with its exact prompts, call the model, persist the assistant turn with the raw payload, mark the step complete — the entire sequence wrapped in a database transaction so a failure never leaves a half-recorded conversation.

Prompt layer. Persona and guardrails at the module level, teaching context at the module level, guidance and model-facing phrasing at the question level, all database-owned with a configuration fallback.

Data layer. A relational schema covering modules, the question tree, question and answer types, answer options, learner answers and their per-type detail records, conversation turns with full prompt and payload storage, and per-question follow-up allowances — with long-text columns for prompt and response content and composite indexing on the conversation lookup path.

Administrative layer. Full CRUD across curriculum, question and answer taxonomy, users and settings, with dedicated prompt editors and rich-text authoring with image upload.

Supporting services. Object storage for uploaded media behind a storage-disk abstraction, error monitoring, an in-application log viewer for administrators, a cache layer with model-level caching on hot read models, and a token-authenticated API for a companion client.

Flow: Learner application → Answer normalisation → Layered prompt composition (persona + module context + question guidance + answer + history) → Rich-text to prompt-text conversion → Language model provider → Transactional persistence of both turns with full prompt and payload → Resumable learner state

Technology Stack

Category Technology
Backend language PHP 8.3
Backend framework Laravel
Database MySQL 8 with migration-managed schema
Cache / queue Redis, with model-level query caching on hot read models
AI integration Hosted large language model chat completions (OpenAI) via an official Laravel client
Web authentication Session authentication with role middleware and CSRF protection
API authentication OAuth2 access tokens (Laravel Passport), Laravel Sanctum
Object storage AWS S3 via Flysystem, behind a storage-disk abstraction
Templating Blade server-rendered views
Frontend Bootstrap 5, Tailwind CSS, Alpine.js, jQuery with AJAX
Asset pipeline Laravel Mix (webpack), Sass, PostCSS, Autoprefixer
Web delivery Progressive Web App support
Monitoring Sentry-compatible error monitoring, in-application log viewer
Testing Pest and PHPUnit
Containerisation Docker — PHP-FPM, nginx, MySQL, Redis and a Node build container
CI / CD Hosted CI pipeline driving scripted release-based deployments
Environments Branch-driven staging, feature and experiment environments alongside production
Code quality Automated style enforcement and static analysis in the pipeline

Technical Challenges & Solutions

Challenge Our Approach
A general-purpose language model answering from general knowledge rather than the course material Layered prompt composition — a module-level persona and guardrail layer, a module-level teaching context layer, a question-level guidance layer, and the learner's answer and history — so every response is grounded in the specific material the question is teaching.
Prompt improvements requiring a developer and a deployment Prompts moved out of code and into the database, edited through dedicated admin screens by the subject-matter experts who understand what a good tutor response looks like, with the original engineered prompts retained in configuration as a fallback.
Explaining a poor AI response after prompts have since been edited Every conversation turn persists the exact system prompt, the exact user prompt and the raw provider payload alongside the message, so what was actually sent is a stored fact rather than a reconstruction.
WYSIWYG-authored teaching content being unusable as prompt input A dedicated normalisation routine that parses the markup properly, decodes entities, normalises whitespace, extracts visible text and rebuilds paragraph, section and list structure — preserving the organisation the meaning depends on rather than flattening it.
Five different answer formats needing to enter one conversational pipeline Every answer type is rendered into a natural-language description of what the learner expressed before the model is called, so the conversation reads coherently regardless of the input widget behind it.
Unbounded conversation cost and latency on a paid third-party API A database-backed follow-up allowance per learner per question, enforced server-side and surfaced in the interface, aligned with the conversational limit the tutoring design itself specifies — plus a step type that completes without any model call where there is nothing to evaluate.
A half-written conversation if a model call or write fails Answer records, per-type detail records and both conversation turns are written inside a single database transaction, so the stored state is always consistent with what the learner actually experienced.
Learners leaving mid-module and returning later Progress, completion, prior answers and full per-question conversation history are restored from the database on load, so the interface returns to exactly the state the learner left.
Iterating on curriculum and prompts without destabilising the main environment Branch-driven environment provisioning in the delivery pipeline, including feature and experiment environments, so a prompt or curriculum variant can be trialled live in isolation.

Security & Reliability

Separated authentication paths. Learner and administrative authentication are distinct, with role middleware enforced on both route groups and administrative access rejected at the middleware layer rather than inside controllers.

An additional gate on prompt management. Curriculum and prompt editing sits behind a further access check beyond ordinary administrative authentication, keeping prompt engineering separate from routine user administration.

Standard web protections. CSRF verification on state-changing requests, password hashing, cookie encryption, request trimming with credential fields excluded, and framework-level validation on every write path.

Token-authenticated API. OAuth2 access tokens govern the programmatic interface, keeping the API surface separate from the session-based web application.

Transactional integrity. The answer-and-conversation write path is transactional, so a failed model call or a failed write rolls back cleanly rather than leaving orphaned records.

Complete AI interaction history. Storing both prompts and the raw provider payload for every turn provides a genuine audit trail of what the platform said to each learner and why — the accountability record any AI-mediated education product needs.

Quota enforcement server-side. Follow-up allowances are enforced on the server and merely reflected in the interface, so the limit is a control rather than a suggestion.

Production monitoring. Error monitoring in production with administrative log inspection for diagnosing issues in the conversation pipeline.

Scalability & Performance

Stateless application tier. Uploaded media is externalised to object storage behind a storage-disk abstraction and session and cache backends are configurable, so application containers hold no local state and scale horizontally.

Cached hot reads. Settings and user records are read on effectively every request and change rarely, making them the right candidates for model-level query caching.

Indexed conversation retrieval. The conversation table carries composite indexing on the exact combination the learner view queries, keeping history retrieval efficient as transcripts accumulate.

Cost-aware model usage. Steps requiring no evaluation never reach the model, and follow-up conversation is quota-bounded — two design decisions that cap per-learner API spend without degrading the experience.

Perceived latency management. Answer submission and conversation run asynchronously with typing indicators, so the learner sees immediate acknowledgement while the model responds rather than a blocked page.

Container-based delivery. A reverse proxy in front of the application process manager, with assets compiled at deploy time and application caches cleared and rebuilt as part of the release.

Release discipline. Multi-release retention with shared linked directories, so a deployment is reversible and uploaded content survives releases.

Business Outcomes

  • Expert-style tutoring is delivered to every learner, on every question, rather than being rationed to whoever can secure time with a senior practitioner.
  • Educators own the quality of the tutoring. Persona, teaching context and question guidance are edited directly in the admin interface, so improving a tutor response is a content change made by the person who spotted the problem.
  • AI behaviour is explainable. Every response can be traced to the exact prompt that produced it, which makes quality review a routine process rather than an investigation.
  • Curriculum authoring is self-service. Modules, sections, steps, answer formats and options are all managed through the back office without developer involvement.
  • Learners can pause and resume without losing anything, including their conversations, which materially affects completion in a working-professional audience.
  • Conversation cost is predictable, through server-enforced allowances and steps that skip the model where nothing needs evaluating.
  • Prompt and curriculum experiments have somewhere safe to run, in isolated environments provisioned from branches.
  • One pipeline serves every question format, so adding a new answer type does not mean building a new tutoring path.

Why it worked

Building on top of a language model is easy for a week and hard for a year. The first version always works: a prompt, an API call, a convincing answer. The difficulty arrives afterwards — when the responses are subtly wrong for reasons nobody can reconstruct, when the only person who can improve them is the engineer who wrote the prompt into the source, and when the cost per learner turns out to be whatever the model felt like generating.

We built this system around the three decisions that determine whether an AI product survives that year. Prompts are content, so the people with the domain expertise can improve them directly and immediately. Every turn stores the exact prompt that produced it, so behaviour is explainable rather than mysterious. And the conversation is bounded by design — quotas enforced server-side, and steps that skip the model entirely where there is nothing to evaluate — so cost is an engineering property rather than a monthly surprise.

The rest is disciplined application engineering: transactional writes so a failed API call never leaves a corrupted conversation, a normalisation layer so authored rich text reaches the model as clean structured prose, one pipeline that absorbs every answer format, and fully resumable state so a learner returning after a fortnight continues rather than restarts.

Our teams work across Laravel and modern PHP, LLM application architecture and prompt systems, content management and authoring workflows, containerised delivery and branch-driven CI/CD — with the judgement to know which parts of an AI product need to be engineered for change and which can be left simple.

Final Summary

The most valuable interaction in professional education is a senior practitioner reading a learner's reasoning and telling them precisely where it goes wrong. It is also the interaction that caps how many learners a course can serve. Automated marking scales but cannot diagnose; expert tutoring diagnoses but cannot scale.

Our team built a platform that reproduces the tutoring conversation at scale. Learners work through structured case-study modules, answering questions in whatever format suits the material — free text, selections, grouped matching, ratings — and then converse with an AI tutor whose responses are grounded in that module's teaching material and that question's specific guidance. Every answer format is normalised into natural language before the model sees it, so a single conversational pipeline serves the whole curriculum. Progress and conversations are fully resumable, and follow-up is bounded by a server-enforced allowance that keeps cost predictable and conversations purposeful.

What makes it maintainable is the prompt architecture. The tutor's persona and guardrails, the module's teaching context and each question's guidance are database-owned content, edited by subject-matter experts through dedicated admin screens rather than by developers through deployments. Every conversation turn stores the exact prompts that generated it and the raw provider response, so any tutor response can be explained months later even after the prompts have moved on. Rich text authored in a WYSIWYG editor is converted to clean, structurally faithful prompt text rather than flattened. And prompt and curriculum experiments run in environments provisioned from branches, so improving the tutor never means risking the platform.

The result is a learning product where the quality of the AI is owned by the educators, the behaviour of the AI is explainable to anyone who asks, and the cost of the AI is a decision rather than an outcome.

01 — Questions

asked about this kind of project

How do you stop a language model answering from general knowledge instead of the course material?

By layering the prompt and giving each layer a job. A system layer defines the persona, the tone and the behavioural guardrails. A module layer supplies the teaching content the answer must be grounded in. A question layer supplies the specific guidance for what this exercise is teaching. The learner's answer and the prior conversation complete the picture. When responses drift, the fix is almost always a missing or weak layer rather than a different model.

Should AI prompts live in code or in the database?

In the database, once the product is past its first version. Prompts are the main quality lever in an AI product, and the people best placed to pull that lever are the domain experts, not the engineers. Keeping prompts in source means every improvement waits for a release and passes through someone who cannot judge whether it is an improvement. Moving them into managed content — with an editing interface and a safe fallback — puts iteration where the expertise is.

How do you debug a bad AI response weeks after it happened?

Store the prompt with the response. Every conversation turn should persist the exact system prompt, the exact user prompt and the raw provider payload alongside the message. Without that, prompts get edited, content gets updated, and the input that produced a given output becomes unrecoverable. It costs some storage and it is the difference between a quality process and guesswork.

How do you control LLM API costs in a learning product?

Two ways, both structural rather than reactive. First, do not call the model where there is nothing to evaluate — acknowledgement steps and navigation should never reach the API. Second, bound the conversation with a server-enforced allowance per learner per question, surfaced in the interface so the limit is visible rather than a surprise. Rate limiting after the fact treats the symptom; designing the call pattern treats the cause.

Why does rich text need special handling before it reaches a model?

Because authored content is markup and models read prose. Passing raw HTML wastes tokens on tags and confuses the structure. Stripping tags naively collapses headings, numbered sections and bullet lists into an undifferentiated paragraph, and the organisation of teaching material is part of its meaning. The right approach parses the content properly, extracts the visible text, and reconstructs the structure as clean prose the model can follow.

How do you support many different question formats with one AI pipeline?

Normalise before you call. Convert every answer type — free text, single selection, grouped selections, image matching, rating scales — into a natural-language description of what the learner actually expressed, then feed that single representation into one conversational pipeline. The alternative, a separate prompt path per format, multiplies the surface you have to maintain and guarantees the formats drift apart in quality.

Can an AI tutor replace human instructors?

It replaces a specific bottleneck, not the instructor. What it delivers at scale is the immediate, grounded, per-question feedback conversation that no course can otherwise offer every learner. Curriculum design, the teaching material and the tutoring guidance all remain human work — and in a well-built system they remain human-editable, because that is where the quality actually comes from.

What should be built first in an AI-assisted product?

The prompt management and audit layer, before the features that sit on top of it. It is the least visible part of the build and the part that determines whether the product can be improved after launch. Teams that hardcode prompts and store only responses ship faster and then spend the following year unable to explain or reliably improve their own AI behaviour.

03 — Similar project?

describe what is different about yours

Build yours.

Tell me what exists, what you need and the deadline. Fixed price for a defined scope, or hourly from $10 with a written estimate first.

Ahmedabad, India · IST (UTC+5:30) · --:-- IST · Mon–Fri 09:00–18:00 IST · US & EU overlap daily

No newsletter, no CRM. Just a reply.