Delivered remotely for a business based in Seattle, United States. Names withheld by agreement.
Project Overview
Industry: Professional education technology — structured, case-study-based training for working practitioners.
Type of solution: A web-based learning platform with an integrated conversational AI tutor, a full curriculum and prompt authoring back office, and a supporting REST API.
Business context: Advanced professional training depends on expert feedback, and expert feedback does not scale. A senior practitioner reviewing a learner's reasoning one-to-one is the highest-value part of any professional course and the part that caps enrolment. Conventional e-learning solves this by removing the interaction: automated marking can tell a learner they are wrong but cannot tell them why, cannot probe the reasoning behind an answer, and cannot adapt when the misunderstanding is different from the one the question anticipated.
General users: Learners working through course modules, subject-matter experts authoring curriculum and tutoring prompts, and platform administrators.
General purpose: To deliver structured, guided professional training in which every learner receives a substantive, context-grounded conversation about their own reasoning — and to put control of that conversation's quality in the hands of the educators rather than the engineers.
The Business Challenge
Expert feedback is the product, and it does not scale. The instructional value sits in a 1:1 conversation with someone who knows the subject. Those people are scarce, expensive, and busy. Any platform that cannot reproduce that conversation is selling a lesser product.
Automated marking cannot diagnose reasoning. A multiple-choice engine knows whether an answer matches a key. It does not know that the learner reached the right answer for the wrong reason, or the wrong answer through sound logic with one faulty assumption — which is the feedback that actually changes how someone thinks.
A general-purpose model is not a subject-matter tutor. An unguided language model will answer from general knowledge, in a generic voice, at whatever length it chooses. For a professional course, responses have to be grounded in this module's material, aligned with this question's teaching intent, delivered in a consistent voice, and bounded in length and scope.
The people who can fix the tutor are not engineers. When a tutor response is unhelpful, the correction is almost always a prompt change — more context, a sharper guardrail, a clearer instruction. If every prompt change requires a developer and a deployment, iteration slows to the speed of the release cycle and the educators lose ownership of the thing they are best placed to improve.
AI behaviour has to be explainable after the fact. When a learner reports a poor tutor response, the team needs to know exactly what was sent to the model at that moment. If prompts have since been edited, and only the response was stored, that question is unanswerable.
Answers come in many shapes but the conversation has one. Text responses, single and grouped selections, image-based matching exercises and rating scales all have to enter the same conversational pipeline and read as something a person actually said.
Every model call costs money and time. Unbounded follow-up chat is unbounded expense, and a third-party API call inside a page interaction is latency the learner feels.
Learning is not one sitting. Learners leave mid-module and come back. Progress, answers and the full conversation have to be exactly where they left them.
Our Approach
Make the prompts editable content, owned by educators. The tutor's persona and guardrails live on the module. The module's teaching material lives on the module. Each question carries its own guidance and its own model-facing phrasing. All of it is stored in the database and edited through purpose-built prompt editors in the admin interface, separate from the ordinary content forms and behind an additional access gate. Improving the tutor is a content change, not a release.
Compose the prompt in layers with clear ownership. Persona and behavioural guardrails form the system layer. The module context, the question's guidance, the learner's answer and the prior conversation form a structured, labelled user layer. Each layer has one owner and one editing surface, which keeps prompt engineering tractable as the curriculum grows.
Store the prompt with the response, every turn. Each conversation turn persists its role, its content, the exact system prompt and user prompt used to generate it, and the raw provider payload. When a tutor response needs to be explained, the team reads exactly what was sent — not an approximation reconstructed from prompts that have since changed.
Treat rich text as a prompt engineering problem. Teaching material is authored in a WYSIWYG editor and stored as markup. Sending that markup to a model wastes tokens and degrades output; stripping tags naively destroys the structure the meaning lives in. We built a dedicated normalisation routine that parses the content properly, decodes entities, normalises whitespace, extracts the visible text, and then reconstructs paragraph, section and list structure so the model receives clean, well-formed prose with its organisation intact.
Normalise every answer type into natural language before the model sees it. Selections, groupings, ratings and free text are all rendered into a sentence describing what the learner actually expressed, so the conversation reads coherently no matter which input widget produced it.
Bound the conversation deliberately. A per-learner, per-question follow-up allowance is tracked in the database and reflected in the interface, so the technical limit and the tutor's own conversational limit agree. Where a step requires only acknowledgement and there is nothing to evaluate, the model is not called at all.
Make everything resumable. Completed questions, prior answers, section-level and overall progress, and the full conversation for every question are rehydrated from the database on load, so a learner returning after a week continues rather than restarts.
Give prompt iteration somewhere to run. The delivery pipeline creates environments from branches — including per-feature and experiment environments — so a curriculum or prompt variant can be trialled in a live environment without disturbing the main staging line.
The Solution
Module and curriculum authoring. Modules are composed of sections, each containing an ordered sequence of learner steps, with drag-and-drop reordering, rich-text content, image upload and presentation types controlling how each step is rendered.
Layered prompt management. Each module carries a tutor persona prompt and a body of teaching context. Each question carries its own guidance and its own model-facing phrasing. Both are edited through dedicated prompt editors, with the original engineered prompts retained in configuration as a fallback for modules not yet given their own.
Multiple answer formats. Free text, single selection, grouped selection, image-based grouped matching, rating scales, and an acknowledgement step that completes without evaluation — each with its own authoring form and its own rendering.
Grounded tutoring conversation. When a learner submits an answer, the platform assembles the persona, the module context, the question guidance, the learner's answer rendered as natural language and the prior conversation, and returns a response that corrects misunderstandings, probes reasoning or asks a clarifying question.
Bounded follow-up. Learners can continue the conversation within a question up to a tracked allowance, with the remaining count surfaced in the interface and enforced server-side.
Full conversation audit trail. Every turn is stored with the exact prompts that produced it and the raw provider response, making tutor behaviour reviewable and diagnosable.
Progress and resumption. Section-level and overall progress indicators, a step counter, completion tracking, and full restoration of answers and conversation on return.
Answer reset. A learner can reset a step, clearing the answer, its detail records and its conversation, and restoring the follow-up allowance.
Administrative back office. Role-separated administration covering modules, questions, question and answer types, answer options, learners, administrators, branding and platform settings, with prompt and curriculum management behind an additional gate.
Supporting API. A token-authenticated REST interface covering registration, authentication, password recovery and profile management for a companion client.
Installable web experience. Progressive Web App support so the learner application can be installed and launched like a native application.
Key Features
Database-managed, layered AI prompts Persona, module context and per-question guidance are stored as content and edited through dedicated admin screens, so the people who understand the subject matter control the quality of the tutoring without a code change.
Full prompt-and-response audit trail Every conversation turn persists the exact system prompt, the exact user prompt and the raw provider payload alongside the message itself, so any tutor response can be explained after the fact even after prompts have changed.
Rich-text to prompt-text normalisation A dedicated routine converts WYSIWYG-authored teaching content into clean, structurally faithful plain text for the model — preserving numbered sections, labelled headings and lists rather than flattening them.
Answer-type normalisation into a single conversational pipeline Selections, groupings, image matching, ratings and free text are all rendered into natural language before reaching the model, so one pipeline serves every question format.
Bounded, quota-managed conversation A per-learner, per-question follow-up allowance is enforced server-side and reflected in the interface, keeping cost predictable and conversations purposeful.
Evaluation-free steps that skip the model entirely Acknowledgement steps complete without an API call, because there is nothing to assess — cost control by design rather than by throttling.
Resumable learning state Progress, answers and complete per-question conversations are restored from the database on load, so learners continue exactly where they stopped.
Two-level curriculum structure with ordered steps Sections and their ordered child steps drive both the sidebar navigation and the progress model, with reordering handled from the admin interface.
Separated prompt-engineering access Curriculum and prompt management sits behind an additional gate beyond ordinary administrative access, distinguishing prompt engineering from routine user administration.
Branch-driven environments for curriculum experiments The delivery pipeline provisions environments from branches, including feature and experiment lines, so prompt and curriculum variants can be trialled live without disturbing the main staging environment.
Technical Architecture
Learner application. Server-rendered templates with a sectioned sidebar, step navigation and progress indicators, enhanced with asynchronous requests for answer submission and conversation so the page never reloads mid-exercise. Typing indicators manage perceived latency while the model responds. Full state is rendered from the database on load, making the experience resumable without client-side persistence.
Application layer. A framework monolith with three controller trees — learner, administrative and API — over shared base controllers. Role middleware separates learner and administrative routes; an additional session-check middleware gates the curriculum and prompt management area.
Conversation pipeline. The core flow is deliberately linear and lives in one place: normalise the submitted answer into natural language, load the module persona and context and the question guidance, convert all rich-text content to clean prompt text, compose the layered user prompt with prior history, persist the learner turn with its exact prompts, call the model, persist the assistant turn with the raw payload, mark the step complete — the entire sequence wrapped in a database transaction so a failure never leaves a half-recorded conversation.
Prompt layer. Persona and guardrails at the module level, teaching context at the module level, guidance and model-facing phrasing at the question level, all database-owned with a configuration fallback.
Data layer. A relational schema covering modules, the question tree, question and answer types, answer options, learner answers and their per-type detail records, conversation turns with full prompt and payload storage, and per-question follow-up allowances — with long-text columns for prompt and response content and composite indexing on the conversation lookup path.
Administrative layer. Full CRUD across curriculum, question and answer taxonomy, users and settings, with dedicated prompt editors and rich-text authoring with image upload.
Supporting services. Object storage for uploaded media behind a storage-disk abstraction, error monitoring, an in-application log viewer for administrators, a cache layer with model-level caching on hot read models, and a token-authenticated API for a companion client.
Flow: Learner application → Answer normalisation → Layered prompt composition (persona + module context + question guidance + answer + history) → Rich-text to prompt-text conversion → Language model provider → Transactional persistence of both turns with full prompt and payload → Resumable learner state
Technology Stack
| Category | Technology |
|---|---|
| Backend language | PHP 8.3 |
| Backend framework | Laravel |
| Database | MySQL 8 with migration-managed schema |
| Cache / queue | Redis, with model-level query caching on hot read models |
| AI integration | Hosted large language model chat completions (OpenAI) via an official Laravel client |
| Web authentication | Session authentication with role middleware and CSRF protection |
| API authentication | OAuth2 access tokens (Laravel Passport), Laravel Sanctum |
| Object storage | AWS S3 via Flysystem, behind a storage-disk abstraction |
| Templating | Blade server-rendered views |
| Frontend | Bootstrap 5, Tailwind CSS, Alpine.js, jQuery with AJAX |
| Asset pipeline | Laravel Mix (webpack), Sass, PostCSS, Autoprefixer |
| Web delivery | Progressive Web App support |
| Monitoring | Sentry-compatible error monitoring, in-application log viewer |
| Testing | Pest and PHPUnit |
| Containerisation | Docker — PHP-FPM, nginx, MySQL, Redis and a Node build container |
| CI / CD | Hosted CI pipeline driving scripted release-based deployments |
| Environments | Branch-driven staging, feature and experiment environments alongside production |
| Code quality | Automated style enforcement and static analysis in the pipeline |
Technical Challenges & Solutions
| Challenge | Our Approach |
|---|---|
| A general-purpose language model answering from general knowledge rather than the course material | Layered prompt composition — a module-level persona and guardrail layer, a module-level teaching context layer, a question-level guidance layer, and the learner's answer and history — so every response is grounded in the specific material the question is teaching. |
| Prompt improvements requiring a developer and a deployment | Prompts moved out of code and into the database, edited through dedicated admin screens by the subject-matter experts who understand what a good tutor response looks like, with the original engineered prompts retained in configuration as a fallback. |
| Explaining a poor AI response after prompts have since been edited | Every conversation turn persists the exact system prompt, the exact user prompt and the raw provider payload alongside the message, so what was actually sent is a stored fact rather than a reconstruction. |
| WYSIWYG-authored teaching content being unusable as prompt input | A dedicated normalisation routine that parses the markup properly, decodes entities, normalises whitespace, extracts visible text and rebuilds paragraph, section and list structure — preserving the organisation the meaning depends on rather than flattening it. |
| Five different answer formats needing to enter one conversational pipeline | Every answer type is rendered into a natural-language description of what the learner expressed before the model is called, so the conversation reads coherently regardless of the input widget behind it. |
| Unbounded conversation cost and latency on a paid third-party API | A database-backed follow-up allowance per learner per question, enforced server-side and surfaced in the interface, aligned with the conversational limit the tutoring design itself specifies — plus a step type that completes without any model call where there is nothing to evaluate. |
| A half-written conversation if a model call or write fails | Answer records, per-type detail records and both conversation turns are written inside a single database transaction, so the stored state is always consistent with what the learner actually experienced. |
| Learners leaving mid-module and returning later | Progress, completion, prior answers and full per-question conversation history are restored from the database on load, so the interface returns to exactly the state the learner left. |
| Iterating on curriculum and prompts without destabilising the main environment | Branch-driven environment provisioning in the delivery pipeline, including feature and experiment environments, so a prompt or curriculum variant can be trialled live in isolation. |
Security & Reliability
Separated authentication paths. Learner and administrative authentication are distinct, with role middleware enforced on both route groups and administrative access rejected at the middleware layer rather than inside controllers.
An additional gate on prompt management. Curriculum and prompt editing sits behind a further access check beyond ordinary administrative authentication, keeping prompt engineering separate from routine user administration.
Standard web protections. CSRF verification on state-changing requests, password hashing, cookie encryption, request trimming with credential fields excluded, and framework-level validation on every write path.
Token-authenticated API. OAuth2 access tokens govern the programmatic interface, keeping the API surface separate from the session-based web application.
Transactional integrity. The answer-and-conversation write path is transactional, so a failed model call or a failed write rolls back cleanly rather than leaving orphaned records.
Complete AI interaction history. Storing both prompts and the raw provider payload for every turn provides a genuine audit trail of what the platform said to each learner and why — the accountability record any AI-mediated education product needs.
Quota enforcement server-side. Follow-up allowances are enforced on the server and merely reflected in the interface, so the limit is a control rather than a suggestion.
Production monitoring. Error monitoring in production with administrative log inspection for diagnosing issues in the conversation pipeline.
Scalability & Performance
Stateless application tier. Uploaded media is externalised to object storage behind a storage-disk abstraction and session and cache backends are configurable, so application containers hold no local state and scale horizontally.
Cached hot reads. Settings and user records are read on effectively every request and change rarely, making them the right candidates for model-level query caching.
Indexed conversation retrieval. The conversation table carries composite indexing on the exact combination the learner view queries, keeping history retrieval efficient as transcripts accumulate.
Cost-aware model usage. Steps requiring no evaluation never reach the model, and follow-up conversation is quota-bounded — two design decisions that cap per-learner API spend without degrading the experience.
Perceived latency management. Answer submission and conversation run asynchronously with typing indicators, so the learner sees immediate acknowledgement while the model responds rather than a blocked page.
Container-based delivery. A reverse proxy in front of the application process manager, with assets compiled at deploy time and application caches cleared and rebuilt as part of the release.
Release discipline. Multi-release retention with shared linked directories, so a deployment is reversible and uploaded content survives releases.
Business Outcomes
- Expert-style tutoring is delivered to every learner, on every question, rather than being rationed to whoever can secure time with a senior practitioner.
- Educators own the quality of the tutoring. Persona, teaching context and question guidance are edited directly in the admin interface, so improving a tutor response is a content change made by the person who spotted the problem.
- AI behaviour is explainable. Every response can be traced to the exact prompt that produced it, which makes quality review a routine process rather than an investigation.
- Curriculum authoring is self-service. Modules, sections, steps, answer formats and options are all managed through the back office without developer involvement.
- Learners can pause and resume without losing anything, including their conversations, which materially affects completion in a working-professional audience.
- Conversation cost is predictable, through server-enforced allowances and steps that skip the model where nothing needs evaluating.
- Prompt and curriculum experiments have somewhere safe to run, in isolated environments provisioned from branches.
- One pipeline serves every question format, so adding a new answer type does not mean building a new tutoring path.
Why it worked
Building on top of a language model is easy for a week and hard for a year. The first version always works: a prompt, an API call, a convincing answer. The difficulty arrives afterwards — when the responses are subtly wrong for reasons nobody can reconstruct, when the only person who can improve them is the engineer who wrote the prompt into the source, and when the cost per learner turns out to be whatever the model felt like generating.
We built this system around the three decisions that determine whether an AI product survives that year. Prompts are content, so the people with the domain expertise can improve them directly and immediately. Every turn stores the exact prompt that produced it, so behaviour is explainable rather than mysterious. And the conversation is bounded by design — quotas enforced server-side, and steps that skip the model entirely where there is nothing to evaluate — so cost is an engineering property rather than a monthly surprise.
The rest is disciplined application engineering: transactional writes so a failed API call never leaves a corrupted conversation, a normalisation layer so authored rich text reaches the model as clean structured prose, one pipeline that absorbs every answer format, and fully resumable state so a learner returning after a fortnight continues rather than restarts.
Our teams work across Laravel and modern PHP, LLM application architecture and prompt systems, content management and authoring workflows, containerised delivery and branch-driven CI/CD — with the judgement to know which parts of an AI product need to be engineered for change and which can be left simple.
Final Summary
The most valuable interaction in professional education is a senior practitioner reading a learner's reasoning and telling them precisely where it goes wrong. It is also the interaction that caps how many learners a course can serve. Automated marking scales but cannot diagnose; expert tutoring diagnoses but cannot scale.
Our team built a platform that reproduces the tutoring conversation at scale. Learners work through structured case-study modules, answering questions in whatever format suits the material — free text, selections, grouped matching, ratings — and then converse with an AI tutor whose responses are grounded in that module's teaching material and that question's specific guidance. Every answer format is normalised into natural language before the model sees it, so a single conversational pipeline serves the whole curriculum. Progress and conversations are fully resumable, and follow-up is bounded by a server-enforced allowance that keeps cost predictable and conversations purposeful.
What makes it maintainable is the prompt architecture. The tutor's persona and guardrails, the module's teaching context and each question's guidance are database-owned content, edited by subject-matter experts through dedicated admin screens rather than by developers through deployments. Every conversation turn stores the exact prompts that generated it and the raw provider response, so any tutor response can be explained months later even after the prompts have moved on. Rich text authored in a WYSIWYG editor is converted to clean, structurally faithful prompt text rather than flattened. And prompt and curriculum experiments run in environments provisioned from branches, so improving the tutor never means risking the platform.
The result is a learning product where the quality of the AI is owned by the educators, the behaviour of the AI is explainable to anyone who asks, and the cost of the AI is a decision rather than an outcome.