OFM gamification, chatter academy, and creator engagement domain
Date: 2026-08-29 Owner: OFM-TRAINING-DEEP Status: complete; urgent map accepted; external sourcing, ranking, composition, and proof plan recorded
Executive decision (initial)
HALO should own the training and behavior-control plane: versioned creator voice and boundary playbooks, synthetic fan scenarios, branching simulation state, rubric versions, evidence-linked scores, calibration, coaching, certification, and the authoritative XP / reward ledger. It should also own a separate, opt-in creator/model progress loop for onboarding, content/approval missions, availability, boundaries, and agency goals. Buzz is the collaboration and agent plane: it can host discussion, reminders, assignments, handoffs, and approval conversations, but it must not become the source of truth for scores, certification, creator consent, or rewards.
This is not a generic LMS search. The target outcome is a chatter who can hold a creator-aligned conversation safely and consistently, recognize intent, handle objections, offer PPV/custom work without coercion or misrepresentation, preserve boundaries, escalate when needed, and leave a durable handoff. Gross sales are one signal, never the sole training objective or reward trigger.
Fact from the read-only clone: HALO has a functioning Tasks & Rewards surface with quests, per-player cadence slots, re-rolls, screenshot evidence, pending/verified/rejected completion, manager review, XP, HALO currency, a shop/voucher flow, leaderboard, and shift-scoped quest assignments. Inference: this is a gamification substrate, not yet a chatter academy or a quality-assured behavior system. Product proposal: add training and creator engagement as separate state machines that consume authoritative domain events and preserve the existing rewards surface only where its evidence and safety gates are strong enough.
Persona and authority matrix
The table describes the target operating model. “Current HALO evidence” is deliberately separate from the target authority so existing screens do not define the product boundary.
| Persona / service | Must see | May change | Authoritative record / hard boundary | Current HALO evidence |
|---|---|---|---|---|
| Owner / admin | Agency-wide training health, quality and safety trends, certification risk, reward liability, creator wellbeing signals, audit and appeals | Training policy, reward budget, visibility rules, escalation thresholds, break-glass access, final appeals | HALO policy/config, audit, certification and reward records; cannot rewrite past evidence | Admin can manage quests and shop items; no training control centre or fairness/quality view is evidenced (repo/src/components/gamification/QuestsPanel.tsx, repo/convex/gamification.ts:184-200). |
| Training manager / DCR | Playbook versions, scenario library, cohorts, rubric versions, calibration queue, skill gaps and interventions | Author/submit scenario and rubric drafts, assign cohorts, propose coaching, run calibration, verify bounded completions | HALO playbook/scenario/rubric/version records; publication requires review; cannot silently alter historical scores | DCR is a quest manager and player in the game module (repo/src/hooks/useGameRole.ts:3-16, repo/convex/gamification.ts:184-200), but no scenario/rubric authoring is present. |
| QA reviewer / calibrator | Redacted production samples, simulation transcript, creator boundary snapshot, rubric version, evidence and prior calibration | Score criteria, cite evidence, request escalation, record disagreement and calibration decision | Append-only evaluation and calibration records; no direct change to creator consent, compensation, or raw credentials | HALO has a pending quest-completion review queue, but it reviews attachments rather than conversation quality (repo/convex/gamification.ts:442-467, repo/src/components/gamification/QuestEvidenceUpload.tsx:128-170). |
| Chatter / sales operator | Assigned academy work, synthetic fan state, approved creator playbook version, own transcript/replay, feedback, certification and safe quests | Respond in simulation, request replay/help, acknowledge coaching, submit bounded quest evidence, report a boundary conflict | HALO training run, evidence, skill and certification records; production platform conversation remains platform-authoritative; cannot edit scores or playbook history | Chatter can play quests, upload screenshots, re-roll once per period, earn XP/HALO, and buy vouchers (repo/src/components/gamification/ChatterQuestsPage.tsx:1-280, repo/src/components/gamification/QuestEvidenceUpload.tsx:74-170). |
| Creator / model | Own boundary/voice/availability record, assigned onboarding/content/approval missions, progress, feedback addressed to them, optional recognition and agency goals | Supply/revoke boundaries and consent; set availability; approve how their voice is represented; complete or decline missions; request correction/export | Creator profile and consent/boundary records in HALO; creator autonomy and wellbeing override engagement targets; no staff ranking or private chatter notes | Creator onboarding, profile, calendar, announcements, upload and model data routes exist, but no creator-facing training/gamification surface is proven (repo/src/App.tsx:145-193; claimed separate model app remains open). |
| Playbook / content author | Creator-approved voice notes, boundaries, prohibited topics, offers, escalation rules, policy references | Draft playbook blocks and scenario fixtures; submit for creator/manager review | Versioned playbook with author, source evidence, effective window and approval chain; never infer consent from a chat message | HALO stores model-profile fields and announcements, but no versioned playbook or citation graph is evidenced (repo/src/components/creators-data/CreatorDataModal.tsx:897-1055). |
| HR / workforce / people operations | Attendance and certification status, coaching completion, aggregate fairness/wellbeing metrics | Assign required training, manage employment-side acknowledgement and reward fulfillment | Workforce records remain separate; training data is minimized and purpose-limited; no access to raw fan content by default | Payroll and team routes exist, but no training/certification linkage is shown in the known map (repo/src/App.tsx:150-183). |
| Simulation / evaluator service | Synthetic scenario spec, redacted response, allowed playbook/policy excerpts, model/version and run seed | Generate a fan turn, classify rubric candidates, explain uncertainty, suggest feedback | HALO run/evaluation evidence; provider is untrusted and cannot certify, grant rewards, contact a real fan, or access raw creator secrets | No simulation/evaluator service or scenario tables appear in the known training/gamification surface; this is a campaign finding to verify broadly. |
| HALO domain services | Stable creator/chatter IDs, consent, content/request links, approved revenue/quality events, audit context | Emit or accept typed, idempotent events and bounded commands | Domain services remain authoritative for creators, consent, content, conversations, money, and audit; training consumes references rather than copying truth | Existing gamification functions accept IDs and client-shaped values; the module comments describe caller-supplied profile IDs and no general auth context, requiring hardening before expansion (repo/convex/gamification.ts:4-17). |
| Buzz collaboration / agent plane | Scoped training assignments, summaries, handoff links, reminders, review discussion and escalation status | Coordinate people and bounded tools; acknowledge or route work | Buzz owns communication/event surfaces; HALO owns playbook, score, certification, reward and consent state | Buzz is fixed by the campaign brief; no training authority should be placed there. |
End-to-end operating journey
The authoritative record in each step is the proposed target record. Existing HALO coverage is mapped later and must not be mistaken for implementation authorization.
- Define the outcome and safety envelope. The manager names the skill and acceptable
outcome: e.g. understand a fan’s intent, respond in the creator’s approved voice, offer an approved PPV/custom path, preserve consent and boundaries, and leave a clear follow-up. The system records quality, safety, retention and escalation measures before any gross-sales target. Authority: training policy plus the relevant creator boundary snapshot.
- Resolve the current creator playbook. A creator or authorized manager provides
voice examples, permitted offers, boundaries, availability, prohibited claims, escalation rules and preferred handoff language. A playbook block is versioned, source-linked, effective-dated and explicitly approved. Expiry or revocation blocks affected scenarios and production coaching until a new version is published. Authority: HALO playbook / consent records; Buzz is only a discussion surface.
- Author a synthetic fan and scenario. A manager chooses an archetype (curious new
subscriber, price-sensitive buyer, repeat spender, lapsed fan, custom-request seeker, boundary-testing or angry fan), intent, context, constraints, risk flags, expected branches, difficulty and success conditions. The scenario references exact playbook and policy versions, not a mutable prompt hidden in a vendor. Authority: immutable scenario definition and version.
- Assign a diagnostic or cohort run. The manager assigns a scenario to a chatter or
cohort with due window, attempt limits, accessibility settings and a reason. A baseline run establishes skills without penalizing a first attempt. Buzz can notify and coordinate the shift, but the assignment and due state live in HALO. Failure state: stale playbook, missing boundary, or inaccessible scenario blocks launch and creates an actionable exception.
- Run the role-play. The chatter responds turn by turn to a synthetic fan. Branches
react to the chatter’s language, offer choice, consent handling, escalation and follow-up. The chatter can pause, ask for a hint, replay a turn, or escalate. AI may generate a candidate turn but must identify model/version, confidence and scenario seed. No message is sent to a real fan and no production credential is needed.
- Capture evidence and score. Store the transcript, turn IDs, scenario branch,
playbook/policy citations, model/evaluator versions, response timestamps and explicit human actions. Rubric dimensions should include creator-voice fidelity, boundary and consent adherence, intent discovery, empathy, objection handling, offer relevance, truthfulness, escalation/handoff quality, retention thinking and efficiency. Automated scoring is provisional and confidence-labeled; human review is required for safety, policy conflict, low confidence, appeals and certification thresholds. Authority: immutable run/evidence/evaluation records.
- Review, replay and calibrate. The chatter sees evidence-linked feedback, a better
response, why the response matters, and the exact playbook/policy source. They can replay the failed branch with a new attempt. Reviewers compare anonymized examples in a calibration session, record disagreements and publish a rubric version; past scores are not rewritten. Failure states include unsafe content, hallucinated citation, ambiguous rubric, evaluator disagreement and provider outage.
- Coach and certify. A skill graph turns failed criteria into targeted micro-lessons,
scenario retries, manager coaching tasks and a certification decision with issuer, rubric version, evidence links, issue/expiry dates and scope (creator/account/offer class). Certification can be suspended when a playbook or policy changes. The system must support appeal, remediation and re-certification; a leaderboard cannot override a safety gate.
- Transfer safely to production. Only the approved playbook version and relevant
certification scope are surfaced in the chatter’s production workspace. Actual platform conversations remain authoritative in the messaging platform/domain service. A sampled, redacted production transcript may create a QA review linked to the chatter, creator, policy version and outcome, but training never mutates the conversation or sends a message by itself. The manager can pause a skill or creator/account when quality or wellbeing signals deteriorate.
- Engage chatters without rewarding harm. Quests and team missions should reward
verified practice coverage, quality improvements, respectful escalation, coaching follow-through, accurate handoffs, retention-safe outcomes and peer support. XP/rewards derive from server-authoritative events and rubric thresholds. Gross sales may be a bounded context signal, never a standalone reward event. Controls include daily caps, duplicate detection, minimum quality gates, opt-out from public ranking, team-safe competition, appeal, and a manager view of reward distribution.
- Run a separate creator/model engagement loop. The creator sees a private progress
view for onboarding, boundary review, content/approval missions, availability, request readiness and agency goals. Goals are transparent, optional where appropriate and never tied to coercive streak loss, hidden compensation, or disclosure of internal chatter performance. A creator can approve/revoke voice and boundaries, decline a mission, request support, and export their history. Creator progress events may inform agency planning but must not be silently mixed into chatter rankings.
- Shift handoff, exit and learning. At shift change, the system carries unresolved
scenario state, promised follow-ups, coaching tasks, current playbook version, creator boundary changes and certification scope through a typed handoff. At staff departure, creator departure, account suspension or platform outage, freeze affected rewards and training assumptions, preserve evidence, revoke access, and export a manifest of runs, evaluations, certifications, playbook approvals, rewards and retention decisions.
Feature and sub-app tree
Training + engagement operating system
├── Training control plane
│ ├── creator voice, boundaries, offers, prohibited claims and escalation rules
│ ├── playbook blocks with source links, effective dates, diff and approval workflow
│ ├── policy / safety library and scenario eligibility gates
│ ├── creator consent and revocation impact analysis
│ └── immutable versions, audit, export and retention
├── Chatter academy
│ ├── skill graph, diagnostic baseline and role/creator/account scope
│ ├── course/micro-lesson catalog and manager-authored assignments
│ ├── cohort, due window, accessibility and attempt policy
│ ├── certification, expiry, suspension, appeal and remediation
│ └── shift handoff and training continuity
├── AI role-play scenario lab
│ ├── synthetic fan archetypes, intent, context, risk and constraints
│ ├── branching scenario graph and expected safe/unsafe paths
│ ├── deterministic seed, model/provider/version and evaluator settings
│ ├── turn-by-turn simulator with pause, hint, replay and human takeover
│ ├── transcript redaction, retention and data-minimization controls
│ └── scenario fixtures, regression tests and adversarial safety probes
├── Assessment, QA and calibration
│ ├── rubric dimensions, weights, threshold and rubric version
│ ├── automated provisional scoring with confidence and evidence spans
│ ├── human review queue, disagreement, appeal and second review
│ ├── calibration sessions with anchor examples and inter-rater drift
│ ├── sampled production QA linked to actual platform evidence
│ └── coaching plans, manager notes and measurable follow-through
├── Chatter progression and rewards
│ ├── skill/quality quests, practice streaks and team missions
│ ├── rewards from authoritative events with anti-gaming caps
│ ├── private progress, opt-in recognition and fair team views
│ ├── quality/conversion/retention/safety scorecards, not gross-only ranks
│ ├── reward ledger, appeals, reversals and redemption controls
│ └── wellbeing pulse, workload guardrails and pause/opt-out
├── Creator/model engagement
│ ├── private onboarding and readiness progress
│ ├── boundaries, availability and voice/likeness approval missions
│ ├── content/request/approval tasks and transparent agency goals
│ ├── optional recognition and rewards with no coercive mechanics
│ ├── support/escalation, correction, revoke and export
│ └── creator-facing feedback distinct from internal staff feedback
├── Manager and owner intelligence
│ ├── cohort health, skill gaps, intervention queue and workload
│ ├── quality-before-gross scorecards and fairness/disparate-impact checks
│ ├── calibration drift, evaluator confidence and scenario coverage
│ ├── creator wellbeing/boundary exceptions and consent revocation holds
│ ├── reward distribution, fraud/duplicate signals and appeals
│ └── audit, incident, retention and restore evidence
└── Shared primitives / Buzz adapter
├── stable IDs, typed links, event schema, occurred-at and idempotency key
├── evidence spans, provenance, redaction, retention and export manifest
├── tenant/resource/actor authorization and purpose limitation
├── append-only score/reward/certification history
├── human approval, pause, resume, cancel, retry and dead-letter states
└── Buzz assignments, notifications, summaries and scoped agent tools
Nice-to-have only after the core proof: peer mentoring, cosmetic badges, narrative themes, multi-evaluator ensembles, adaptive difficulty, mobile push, and public team showcases. HALO should not build a generic video-LMS catalog, a real-fan autonomous messaging bot, or a compensation system whose opaque score decides pay.
Current HALO coverage and gaps (read-only map)
| Capability | Evidence in HALO | Covered today | Material gap / implication |
|---|---|---|---|
| Gamification entry and roles | Tasks & Rewards routes are /tasks-rewards, /quests, /supply-depot, and /control-panel; the page uses separate manager/player navigation (repo/src/App.tsx:98-130,180-183; repo/src/pages/TasksRewards.tsx:1-187). | Authenticated game shell; Admin/DCR manager vs player behavior; admin preview is view-only for participation. | No academy home, cohort/certification navigation, creator engagement route, or training exception queue. The current visual/game shell should be a substrate, not the domain model. |
| Quest catalog and cadence | gamificationQuests stores title, description, type, game name, XP/HALO reward, target and active flag (repo/convex/schema.ts:621-634). Daily/weekly/monthly pages and per-player slots are implemented (repo/src/components/gamification/ChatterQuestsPage.tsx:1-280; repo/convex/gamification.ts:516-535,1191-1305). | Manager-authored quest rows, cadence, assignment windows, shift/dept targeting, re-rolls, progress target. | No skill taxonomy, learning objective, scenario graph, difficulty, rubric version, certification scope, or immutable version history. updateQuest mutates a row and deleteQuest removes attached records (repo/convex/gamification.ts:708-799). |
| Chatter practice evidence | Progress rows store one attachment URL per slot; the UI uploads image evidence, allows replacement/removal, then submits completion (repo/convex/schema.ts:667-674; repo/src/components/gamification/QuestEvidenceUpload.tsx:47-170). | Basic evidence capture, progress count, pending/verified/rejected status, resubmission for shift quests. | Evidence is screenshots/URLs, not a turn-level transcript or citation to a creator playbook/policy. No simulator, fan archetype, branch state, feedback span, evaluator confidence, replay or privacy/redaction contract. |
| Manager review | Pending completions are joined with profiles/assignments/quests; manager can approve or reject and verified completions award rewards (repo/convex/gamification.ts:442-467,940-980). | Human review gate for a quest submission; append-only-ish XP/HALO transactions are written on approval. | No rubric, multi-reviewer calibration, disagreement/appeal, feedback artifact, coaching task, quality threshold, or certification decision. Reward approval is not evidence of skill mastery. |
| Chatter progression | Stats expose total XP and HALO balance; ranks and leaderboard are present; leaderboard includes Chatter/DCR profiles and sorts on total XP (repo/convex/schema.ts:601-619; repo/convex/gamification.ts:282-342). | Visible progress, rank ladder, leaderboard, two currencies and transaction history. | No skill vector, quality/conversion/retention/safety decomposition, confidence, recency/decay, fairness check, workload/wellbeing guardrail, ranking opt-out, or gross-sales anti-gaming policy. Total XP is not competence. |
| Rewards / store | Shop items have cost/stock/status; purchases create voucher codes, debit HALO, decrement stock, and managers can redeem (repo/convex/schema.ts:676-709; repo/convex/gamification.ts:1063-1189; repo/src/components/gamification/SupplyDepot.tsx:1-202). | A bounded internal perk economy and redemption status. | No reward budget/approval, fair-value policy, reversal/appeal, tax/payroll boundary, wellbeing safeguard, or proof that items are lawful/appropriate. Reward ledger is not linked to verified quality events beyond quests. |
| Shift operations | Shift assignment/completion tables, three named shifts, player-by-shift overview and manager verification exist (repo/convex/schema.ts:745-772; repo/convex/gamification.ts:566-687,1393-1504). | Shift-scoped quest assignment and completion visibility. | No shift handoff record, unresolved promise transfer, coaching continuity, coverage/SLA, or link to chatter conversation state. |
| Manager authoring | Control Panel imports quest/shop/assignment components and exposes create/update/delete/assign/word-of-day operations (repo/src/components/gamification/QuestsPanel.tsx:35-109; repo/convex/gamification.ts:708-846,1063-1108,1360-1390). | Basic admin CRUD and daily word. | No scenario DSL, archetype library, playbook editor, rubric/version approval, cohort/intervention queue, or evaluation/calibration tool. |
| Trust boundary and reward integrity | Server-side role helper checks Admin/DCR for manager-only operations (repo/convex/gamification.ts:184-200). | Some manager mutations are role-gated; UI mirrors this (repo/src/hooks/useGameRole.ts:40-57). | Several player-facing mutations accept chatterId/URLs/reward amounts as arguments and do not derive the actor or quest reward in the handler (repo/convex/gamification.ts:911-1000,1010-1059; repo/src/components/gamification/QuestEvidenceUpload.tsx:147-154). The canonical repo audit also records zero ctx.auth/getUserIdentity uses in convex/; this must be resolved before training scores or rewards become authoritative. |
| Creator/model surface | Public/tokenized onboarding, contract signing, upload, internal Model Profile, calendar and announcements are routed (repo/src/App.tsx:145-193; repo/src/components/creators-data/CreatorDataModal.tsx:1008-1055). | Creator data intake and some internal model information / planning surfaces. | No evidenced creator academy, private progress/reward view, boundary/voice playbook approval workflow, certification feedback, wellbeing controls, or export/revoke surface. A separate model app remains an open discovery question; this report does not treat the route map alone as proof of absence. |
| Claimed model app | Known routes include /upload/:id, tokenized onboarding and signing, and internal Model Profile. The prior-art note explicitly says this does not prove a complete model-facing app. | Evidence boundary is honest and bounded. | Broad repository/deployment/product search must locate or rule in/out the claimed app before final creator-gap language. |
Immediate HALO training gaps
The highest-risk gap is not cosmetic gamification; it is authority. The current client can shape completion inputs, while the server stores player-supplied identity and reward amounts. The first training proof must therefore establish server-derived actor identity, server-derived reward values, append-only evaluation/reward events, and a human approval path before adding AI role-play or public rankings.
Internal prior art and crossover
Oracle Streaming evidence (bounded prior art): the protected on-ramp proves a one-action creator/model flow that starts a multi-platform live workflow; unified viewer chat and tips enter one cockpit; real events advance goals/overlays; deduplication and Convex persistence preserve state; human-present/account-safety gates and proof-ledger evidence are explicit (research/ofm-domain-campaign/INTERNAL-PRIOR-ART.md, which routes to apps/oracle-streaming/AGENTS.md, CLAUDE.md, and docs/PROVEN-LEDGER.html).
What transfers:
- a low-cognitive-load “one next action” experience for a chatter’s next training step,
creator approval, shift handoff, and manager review;
- a typed event vocabulary with stable IDs, timestamps, source, deduplication and
idempotency, so goals and rewards advance from real approved events rather than UI counters;
- human-present and account-safety gates as a model for human review, escalation and
certification, not as an excuse to automate creator impersonation;
- proof-ledger discipline: each score, reward and certification state needs evidence a
manager can inspect and replay;
- creator-visible progress that is transparent and separate from internal staff control.
What does not transfer: OBS/OME/browser/readback, live webcam setup, platform-specific chat/tip ingestion, streaming overlays, and live account transport. A streaming tip is not an OnlyFans message, a training pass, or consent to imitate a creator. A shared Convex vendor does not imply a shared schema, auth boundary or deployment.
HALO crossover hypothesis: HALO’s current quest ledger and Oracle’s truth-grounded goal events could converge on a small cross-product event contract—actor_id, creator_id, event_id, event_type, occurred_at, source, evidence_ref, policy_version, idempotency_key—while each product retains its own authority. The training lane should consume approved conversation/QA/playbook events and emit training_run.completed, evaluation.verified, certification.issued, and reward.eligible events. This is an inference and proof target, not an implementation decision.
Claimed model app status: the inspected HALO routes show onboarding, signing, upload and internal model data, but no complete model-facing app can be asserted yet. The external campaign must search the repository/deployments and public/private product evidence before calling creator gamification absent.
First-principles missingness map
The department exists to make human behavior reliably safe, creator-aligned and effective under shift pressure—not merely to make people click more quests. The hidden work between today’s screens is where most value and risk sit.
| First-principles outcome / invisible work | Missing or weak system | Failure if left implicit | 10× proof target |
|---|---|---|---|
| A chatter knows what “good” means for this creator, offer and risk context | Versioned creator voice/boundary/playbook registry with approvals and citations | Generic scripts, boundary drift, unsafe claims, creator mistrust | Every scenario turn cites the active playbook/policy block and blocks on expiry/revocation |
| Training changes behavior, not screenshot submission | Scenario state machine with synthetic fan archetype, intent, branch and replay | Learners optimize evidence upload or memorized answers | Pre/post blinded scenario eval shows improvement on safety, empathy, intent and handoff |
| Managers judge consistently | Rubric versions, anchor examples, calibration, confidence and appeal | Arbitrary scores, favoritism, noisy certification and disputes | Inter-rater agreement/drift is visible; disputed cases get second review without rewriting evidence |
| Real work feeds learning without exposing unnecessary private data | Redacted production QA sample linked to platform evidence, creator, skill and policy | Feedback lives in chat, cannot be replayed, and leaks raw fan/creator content | A sampled conversation becomes an evidence-linked coaching task with retention controls |
| Shift changes do not erase promises or context | Durable handoff object carrying unresolved state, follow-up, current playbook and coaching | Duplicate promises, forgotten fans, repeated mistakes | Incoming chatter can resume from a typed handoff with clear ownership and provenance |
| Rewards drive safe quality rather than raw volume | Server-authoritative event eligibility, quality gates, caps, anti-duplicate rules and appeals | Spam, fake screenshots, coercive pressure, leaderboard gaming | Synthetic adversarial tests cannot mint XP/HALO or certification by client-crafted payloads |
| Creators retain agency while engaging | Private creator surface for boundaries, availability, missions, rewards, pause, revoke and export | Streak pressure, hidden expectation, consent ambiguity and burnout | Creator can see why a goal exists, decline safely, revoke scope and export history |
| Certification remains true after policy/model change | Scope, issuer, rubric/playbook version, expiry and re-certification state | Staff appears certified against stale or revoked guidance | A playbook revocation automatically identifies affected certs and production surfaces |
| AI remains an assistant, not an unreviewed operator | Provider/model/version provenance, redaction, confidence, deterministic seed and human gate | Hallucinated policies, inconsistent scoring, privacy leakage, unreviewed impersonation | Kill-switch, replay and human takeover work during evaluator outage or low confidence |
| Management sees harm as well as revenue | Quality/safety/retention/fairness/wellbeing measures with workload context | Gross sales wins hide coercion, burnout, creator boundary violations or churn | Owner view explains tradeoffs and blocks a reward when safety minimum fails |
| Staff and creators can leave without losing trust evidence | Export manifest for runs, evaluations, certs, playbook approvals, rewards, retention and revocation | Evidence stranded in Buzz/vendor, disputes become unprovable | Synthetic offboarding export round-trips and revokes access without deleting required audit |
| Cross-domain events remain authoritative | Stable IDs and idempotent adapter between HALO, Buzz, messaging, creator and money domains | Duplicate rewards, split-brain status, silent event loss | Replay/out-of-order/duplicate event tests converge to one explainable state |
Initial search vocabulary and coverage receipt
Search intent must be broader than “OnlyFans gamification” or “chatter CRM.” The external campaign will search these families and record negative results as coverage, not silently discard them.
- Conversational sales simulation: conversational sales training, AI role-play,
objection-handling simulator, branching dialogue training, synthetic customer persona, buyer-simulation, sales-assist coaching, call roleplay, chat coaching, fan persona.
- Contact-centre QA/coaching: contact center quality management, conversation QA,
speech/text analytics, agent scorecard, rubric calibration, interaction review, supervisor coaching, side-by-side coaching, workforce learning, escalation quality, customer effort and retention QA.
- Academy / certification: sales enablement LMS, skills graph, competency management,
cohort assignment, diagnostic assessment, learning paths, certification expiry, remediation, evidence-based assessment, rubric engine, competency framework.
- Playbook / knowledge authority: versioned sales playbook, creator voice guide,
policy retrieval, approved response library, source-cited RAG, effective-dated policy, consent-aware knowledge base, boundary management, human approval.
- Gamification / behavioral design: quality-based sales gamification, safe sales
contests, skill progression, team missions, quest engine, habit formation, streak controls, intrinsic motivation, recognition, reward ledger, anti-gaming, fairness, wellbeing, opt-out leaderboard.
- Creator / talent engagement: creator portal progress, model onboarding missions,
talent engagement, influencer operations, creator availability, content approval goals, model wellbeing, creator rewards, rights/consent progress, agency creator community.
- Adjacent industries: healthcare conversational simulation, hospitality service
training, contact-center compliance, financial-services objection handling, insurance claims empathy training, aviation CRM, esports coaching, sports skill progression, marketplace seller education, retail clienteling.
- GitHub intent queries:
branching dialogue simulator,sales roleplay, `synthetic
customer, LMS competency, rubric evaluation engine, human feedback workflow, contact center QA, quest progression, achievement service, streak anti gaming, event sourced rewards, certification workflow, calibration dashboard, skills graph, creator portal gamification`.
- Operator language to validate: chatter training, chatter academy, fan handling,
spender/whale segmentation, PPV unlock, customs, rebill/retention, shift handoff, locked message, girlfriend experience boundaries, upsell without pressure, “script fatigue,” “dead chat,” “lost context,” manager QA, screenshot proof, leaderboard pressure, creator says no, and “what can I promise this fan?” Search terms must be treated as operator vocabulary, not as permission to copy unsafe or exploitative practice.
Urgent coverage receipt: before external sourcing, I read BRIEF.md, ASSIGNMENTS.md, and INTERNAL-PRIOR-ART.md completely; inspected the read-only HALO route map, Tasks & Rewards shell, quest/player/manager/evidence components, gamification schema and Convex queries/mutations, creator onboarding/upload/model-profile routes, and research/CANONICAL-NUMBERS.md. The Serena bridge was unavailable (ECONNREFUSED), so code evidence uses narrow fallback reads only. Local-corpus, GitHub, live-product and operator searches remain pending; no candidate is promoted from this urgent map.
Research status and evidence standard
The urgent map above was written and sent to KELLMAN-SOL before the deep campaign. The following sections are the completed evidence pass. A rank means fit for this HALO authority boundary, not vendor quality, market share, or an adoption recommendation.
Three evidence labels are kept separate:
- Fact from source: a feature or repository property visible in the linked source.
- Vendor claim / operator claim: a product, guide, or case-study statement that is
useful for vocabulary and feature discovery, but is not an independent benchmark.
- Inference / proposal: this lane's conclusion about fit, composition, or build order.
Commercial product pages were treated as private/proprietary SaaS evidence. No vendor metric was used as an OFM benchmark, and no claim of data handling, legal compliance, or platform permission was inferred from a badge or marketing page. GitHub evidence below was retrieved through the core gh api repos/OWNER/NAME endpoint on 2026-08-29, then matched to the root license source returned by gh api repos/OWNER/NAME/contents. No promoted GitHub repository has NOASSERTION; any discovery result with that metadata was excluded from the ranked donor set rather than treated as licensed. No repository was cloned, executed, mutated, committed, or deployed.
Local-corpus and HALO coverage receipt
The local-corpus pass queried both read-only catalogs named by the sourcing campaign:
| Corpus | Read-only evidence | What it was useful for | Limitation |
|---|---|---|---|
SISO_Agent_Base/research/repo-catalog/identity/identity.sqlite | 1,358,200 repo_card rows; narrow exact-name and gamification-description queries | Candidate discovery for skills-service, skills-client, rasa, OpenMAIC, label-studio, argilla, moodle, frappe/lms, and gamification engines | Lexical/index evidence only; broad multi-term scan was too expensive and was stopped, so no timeout was counted as a negative finding |
SISO_Research/siso-foundry/pipelines/github/awesome/catalog_full.sqlite | 307,180-repository curated catalog | Cross-placement and adjacent-category discovery | Curated lexical corpus, not an adoption or license verdict; fresh gh api verification is the authority |
Read-only HALO clone repo/ | research/CANONICAL-NUMBERS.md: 56 schema tables, 39 pages, 36 Convex modules, 14 gamification tables, 0 ctx.auth/getUserIdentity references in Convex, 166 collect() calls | Existing surface, authority gaps, and security-sensitive seams | Clone is evidence only; no live mutation or production claim was made |
The exact local signals were not over-ranked: SkillTree/SkillTree-client were the strongest gamification-specific matches; Rasa and OpenMAIC were simulator/course candidates; Argilla, Label Studio, Promptfoo and Opik were evaluation or feedback infrastructure; Moodle and Frappe LMS were broad LMS candidates; ActiDoo, Laravel Gamify, Level Up and the NGA gamification server were primitive reward substrates. The direct API/license ledger below is the promotion gate.
The most important HALO finding is a trust-boundary blocker, not a missing screen. The read-only clone shows quest completion submission accepting a chatterId, attachments, xpEarned, and bananasEarned in the client flow (repo/src/components/gamification/QuestEvidenceUpload.tsx:128-170, repo/convex/gamification.ts:911-980). Manager role checks exist, but the canonical count shows no general Convex identity enforcement. Therefore the existing quest surface cannot be the authority for certification, compensation-adjacent rewards, or creator access until server-derived actor identity, reward eligibility, evidence provenance, and append-only audit are proven. This report does not change or test those mutations.
Ranked commercial products and live systems
The rank is a buy/pilot/sidecar ranking. HALO remains the authority in every row.
| Rank | Product / class | Direct evidence and best transferable job | Fit, boundary, and disposition |
|---|---|---|---|
| 1 | Hyperbound — AI roleplay + real-call scoring | Its practice product ↗ describes dynamic AI buyers, custom scorecards, roleplay and real-call analytics, LMS/CMS connections, competitions and certifications. Its scorecard documentation ↗, updated 2026-08-26, describes criteria used across roleplays and real calls and comparison of practice with live performance. Vendor/product-doc evidence. | Closest visible loop from synthetic practice to production QA. Pilot only after proving text/chat scenario support, redaction, retention, human review, and export. Vendor score is advisory; HALO stores rubric, evidence, coaching, certification, appeal, and reward truth. |
| 2 | Yoodli — configurable roleplay + programs/certification | Roleplay builder ↗ exposes persona name, voice, tone, demeanor, avatar, pre-read material and goals. Custom goals ↗ support imported rubrics/training material; platform certification ↗ sequences roleplay and accreditation. Product-doc evidence. | Strong candidate for creator-specific voice/boundary drills and structured certification. Voice/video orientation may require a chat adapter. Pilot with synthetic fan text and human calibration; do not let an opaque vendor score decide pay, access, or termination. |
| 3 | Second Nature — AI roleplay academy | Its product page ↗ describes scenario/course authoring, personas, practice, evaluation, progress and knowledge-gap views; its certification article ↗ covers coaching, onboarding and certification. Vendor claim, with visible product workflow. | Good academy/orchestration candidate and useful benchmark for scenario authoring. Public evidence is sales-oriented and does not prove adult chat, consent, or creator voice. Buy/pilot as a bounded simulator; HALO owns source-of-truth playbooks and certification. |
| 4 | Mindtickle — sales enablement + AI roleplay | AI sales role-play ↗ describes realistic scenarios, an AI buyer, objection practice, immediate feedback, roleplay authoring and reviewing multiple submissions; its interactive demo ↗ exposes industry/persona/objection flows. Vendor/product-demo evidence. | Broad enablement system with relevant practice mechanics, but its revenue-enablement center of gravity and closed platform create a higher integration and data-boundary burden. Pilot only for synthetic scenarios; do not delegate OFM policy or certification authority. |
| 5 | MaestroQA — conversation QA + coaching | Its documentation hub ↗ exposes rubric, QA workflow, agent/grader QA, calibration, AI-platform and coaching categories. Its coaching product ↗ describes conversation insights, templates, follow-ups tied to interactions/outcomes and coaching consistency. Product-doc evidence. | Best QA/coaching sidecar among the screened products, not a full simulator. Useful only after HALO defines the sample, redaction, rubric version and appeal contract. Production score remains HALO evidence, not a vendor score. |
| 6 | Playvox — QA/WFM/coaching/gamification | Quality management/coaching material ↗ and its integrated call-center page ↗ describe scorecards, sample distribution, calibration, coaching, badges and leaderboards. Vendor/product-doc evidence. | Valuable for calibration and QA workflow patterns; broad WFM and sales metrics make it unsafe as a single behavioral authority. Consider a QA pilot after privacy and score portability checks. No gross-only contest. |
| 7 | Centrical — adaptive microlearning + performance gamification | Its platform page ↗ describes “signals → intelligence → orchestration → activation → outcomes”; performance management ↗ describes microlearning/coaching signals, points, badges, leaderboards, challenges, redeemable coins and wellbeing. Vendor claim. | Strong pattern for reinforcement and wellbeing-aware engagement, but a black-box commercial control plane is not suitable for HALO certification or rewards authority. Study/pilot a narrow reinforcement feed only. |
| 8 | Axonify — frontline microlearning/reinforcement | Platform material ↗ describes microlearning, guided execution, insights/action, personalized reinforcement, AI content and gamification. Vendor claim. | Good spaced-practice reference for shift workers, weak direct evidence for creator-specific chat QA, appeals and handoff. Study or integrate only as optional learning delivery; HALO owns completion and skill truth. |
| 9 | Docebo — LMS/skills/gamification | Gamified learning ↗ describes challenges, badges, milestones and skills analytics; its developer portal ↗ exposes APIs/webhooks and embedded-learning material; official help ↗ documents badges, leaderboards and rewards-shop concepts. Product-doc evidence. | Viable if agency compliance/course administration becomes a first-order need, but a wholesale LMS would duplicate HALO identity, playbooks and certification. Do not buy before synthetic proof establishes the missing authority contract. |
| 10 | 360Learning / WorkRamp — collaborative LMS/analytics | 360Learning ↗ covers collaborative learning, AI content, skills gaps, onboarding, compliance/re-certification and analytics. WorkRamp reporting ↗ covers dashboards, custom reports, AI analytics and integrations. Product/Help-center evidence. | Useful market comparators for course/catalog/reporting needs, but neither public surface proves synthetic fan simulation, creator voice, or safe chat QA. Study, do not replace HALO with an LMS. |
OFM-adjacent market systems: source events, not training authority
The screened OFM products fill inbox, attribution, staffing, fan-context, or automation jobs. They are useful market evidence and possible lawful event sources, not academy, rubric, or certification systems:
| Product | Evidence | Domain verdict |
|---|---|---|
| ModelVI chatter-management ↗ / CRM ↗ | Vendor claim: per-chatter credentials, shift staffing, sale attribution, conversation QA, fan spend/history/segments and shift continuity. | Potential production-context adapter if terms, creator consent, credential handling and export are proven. It must not issue HALO certifications or rewards. |
| DirtyDialogues ↗ and chat tools ↗ | Vendor/help-center evidence: unified inbox, team management, message attribution, spend/buy-rate context and labels. | Inbox/context donor only; retain raw adult content in the authorized platform/security boundary, send redacted evidence references to HALO. |
| Infloww messages ↗ and fan insights ↗ | Help-center evidence: multi-creator inbox, focus/notifications, spend and creator-note context. | Operational source candidate, not an academy. No training authority or autonomous response delegation. |
| Supercreator ↗ | Vendor claim: AI chatter/CRM/team analytics and routine-chat automation. | Reject as the training authority and as a default integration: autonomous or impersonating responses create consent, policy, and human-approval risk. If studied, use only in a sandbox and require an explicit human gate. |
Market inference: the screened market splits into OFM inbox/attribution, generic simulation, QA/coaching, LMS/reinforcement, and reward layers. No screened product proves the complete combination of creator-approved voice/boundaries, synthetic fan practice, redacted production QA, fair certification, creator consent, and authoritative rewards. That missing combination is why HALO should compose sidecars around a first-party authority plane rather than outsource the domain model.
Operator and market signals
The operator pass validates the shape of the journey, not universal performance standards. The following SOP library ↗, agency-operations guide ↗, trial-shift guide ↗, and skills/rubric guide ↗ repeat the vocabulary of creator-specific voice, agency SOP versus per-creator playbook, scenario/trial-shift grading, fan segmentation, objection/PPV handling, weekly QA, escalation, retention, confidentiality, and shift handoff. These are vendor/operator-guide claims and should become interview prompts and fixture labels, not copied policy.
The training-expectations guide ↗ also names sandbox scenarios, train → practice → QA, appeal, safety and wellbeing. The investigative account ↗ is adversarial context: it supports the existence of training/testing, shadowing and live shift practices while also showing why exploitative scripts and incentive pressure cannot be treated as normative product requirements. Direct OFM interviews remain open; no vendor metric was promoted as a benchmark.
An adjacent trust-and-safety control-loop source, the X DSA transparency report ↗, is useful only for the general diagnose → practice → observe → audit → coach → refresh shape. It is not OFM evidence and does not establish that the same controls or metrics apply to HALO.
Ranked GitHub donors and verification ledger
These are components or patterns to study, adapt behind a HALO boundary, or use in a synthetic proof harness. They are not adoption recommendations. Every cited repository was checked with the core API and its actual root license source; stars and push dates are discovery signals, not quality scores.
| Rank | Repository | Useful donor | Disposition |
|---|---|---|---|
| 1 | NationalSecurityAgency/skills-service ↗ | SkillTree-style micro-learning/gamification service: skills, progress, achievements and learning mechanics. | Study/rebuild the event and progression ideas; Apache-2.0 makes a bounded adapter possible, but Java/service topology is too large to adopt wholesale. |
| 2 | RasaHQ/rasa ↗ | Conversation state/flows and testable dialogue behavior for a synthetic fan simulator. | Study or integrate a sandbox adapter; Apache-2.0. Never treat NLU confidence as a safety or certification decision. |
| 3 | THU-MAIC/OpenMAIC ↗ | Course/session workbench, quizzes, interactives, simulations and reusable learning materials. | Study/prototype; MIT. Fresh release and broad agent surface require security, provenance and persistence review before any use. |
| 4 | argilla-io/argilla ↗ | Human feedback, annotation, continuous evaluation and dataset review. | Integrate as a calibration/evaluator sidecar or study; Apache-2.0. HALO retains rubric and score truth. |
| 5 | HumanSignal/label-studio ↗ | Multi-user labeling and review tied to accounts, with model/active-learning hooks. | Integrate/study for blinded transcript annotation and rubric calibration; Apache-2.0. Not an LMS or certification issuer. |
| 6 | promptfoo/promptfoo ↗ | LLM evaluation, red-team cases, regression suites and CI checks. | Integrate into synthetic evaluator regression; MIT. It proves prompt/test behavior, not human competence or policy compliance. |
| 7 | comet-ml/opik ↗ | Tracing, datasets, LLM-as-judge/evaluation and production monitoring patterns. | Study/integrate only for redacted evaluator telemetry; Apache-2.0. Never use an LLM judge as the sole certification authority. |
| 8 | NationalSecurityAgency/skills-client ↗ | Companion client pattern for SkillTree-style skill/progress presentation. | Study with skills-service; Apache-2.0. Do not copy its client-side progression as authoritative reward state. |
| 9 | moodle/moodle ↗ | Full LMS, competencies, courses and badges. | Study/possible later sidecar; GPL-3.0. Reject wholesale replacement because it duplicates identity, creator boundaries, and HALO authority. |
| 10 | frappe/lms ↗ | Course hierarchy, batches, quizzes, assignments and certificates. | Study/possible later sidecar; AGPL-3.0. License and deployment boundary require legal/architecture review; not a direct chatter simulator. |
| 11 | ActiDoo/gamification-engine ↗ | MIT REST gamification service with achievements, goals, progress, leaderboards, rules, triggers and rewards. | Study/rebuild the minimal event contract; stale relative to other candidates (last push 2023-02-15). Do not delegate reward authority. |
| 12 | cjmellor/level-up ↗ | XP, levels, achievements, streaks, multipliers and experience audit concepts. | Study small patterns; MIT, PHP/Laravel. Use audit/provenance ideas only; streaks must be opt-out and wellbeing-safe. |
| 13 | qcod/laravel-gamify ↗ | Minimal reputation points, badges and user-badge migrations/traits. | Study only; MIT, PHP/Laravel. Too primitive for training evidence, appeal, fairness, or certification. |
| 14 | ngageoint/gamification-server ↗ | Rules/signals, team awards, badges and Open Badges export concepts. | Study only; MIT and stale (last push 2023-07-09). The reward surface is not sufficient evidence for HALO use. |
Core-API verification ledger (captured 2026-08-29)
The values below are the fresh gh api repos/OWNER/NAME fields used for the ranking. The license-source column is the actual root file returned by gh api repos/OWNER/NAME/contents; it is not inferred from a search result. All rows are unarchived. No NOASSERTION row is promoted, so there is no unverified license exception hidden in the ranking.
| Repository | Stars | SPDX from repos | Pushed at | Default branch | License source |
|---|---|---|---|---|---|
NationalSecurityAgency/skills-service | 634 | Apache-2.0 | 2026-08-28T20:58:52Z | master | LICENSE.txt |
RasaHQ/rasa | 21,311 | Apache-2.0 | 2026-07-24T07:39:04Z | 3.6.x | LICENSE.txt |
THU-MAIC/OpenMAIC | 21,275 | MIT | 2026-08-29T07:38:38Z | main | LICENSE |
argilla-io/argilla | 5,088 | Apache-2.0 | 2026-08-24T22:15:29Z | develop | LICENSE |
HumanSignal/label-studio | 28,163 | Apache-2.0 | 2026-08-28T19:16:44Z | develop | LICENSE |
promptfoo/promptfoo | 24,654 | MIT | 2026-08-29T02:53:58Z | main | LICENSE |
comet-ml/opik | 21,662 | Apache-2.0 | 2026-08-28T20:16:26Z | main | LICENSE |
NationalSecurityAgency/skills-client | 103 | Apache-2.0 | 2026-08-26T16:46:07Z | master | LICENSE.txt |
moodle/moodle | 7,362 | GPL-3.0 | 2026-08-18T16:21:56Z | main | COPYING.txt |
frappe/lms | 3,170 | AGPL-3.0 | 2026-08-29T06:39:23Z | develop | license.txt |
ActiDoo/gamification-engine | 474 | MIT | 2023-02-15T22:54:16Z | master | LICENSE |
cjmellor/level-up | 673 | MIT | 2026-07-22T10:00:17Z | 3.x | LICENSE.md |
qcod/laravel-gamify | 679 | MIT | 2026-08-17T07:00:23Z | master | LICENSE.md |
ngageoint/gamification-server | 248 | MIT | 2023-07-09T15:15:15Z | master | LICENSE |
archived=false was returned for each row. The verification method was deliberately
gh api repos/OWNER/NAME --jq '{full_name,html_url,stargazers_count,license:.license.spdx_id,pushed_at,archived,default_branch}'
gh api repos/OWNER/NAME/contents --jq '[.[].name | select(test("(?i)^(license|copying|licence)(\\..*)?$"))]'
The source files were not used to infer a license for any NOASSERTION repository because none was promoted. License compatibility still requires counsel/release review before shipping an adaptation, especially GPL/AGPL components.
Rejected or constrained candidates
| Candidate / pattern | Why it is rejected or constrained for this domain | Safe residual use |
|---|---|---|
| Autonomous OFM chat/PPV automation as training authority | Blurs creator voice, chatter identity, consent, and human approval; a high conversion result can still be unsafe or misleading. | Synthetic sandbox only, or redacted suggestion tooling behind an explicit human gate and HALO policy retrieval. |
| Supercreator-style autonomous chat integration | The public product claim includes routine-chat automation; that is not evidence of creator-approved boundaries, safe escalation, or auditability. | Market/context signal only; no production authority. |
| Gross/revenue-only leaderboard or contest | Rewards volume, unlocks, spend, or raw conversion without safety, consent, retention, creator voice, handoff, and wellbeing gates; easy to game and coercive under shift pressure. | Keep money and attribution in HALO finance/creator domains as one outcome signal; never make it the only score or reward trigger. |
| Spinify/Hoopla-style generic sales gamification as the domain engine | Their public surfaces emphasize CRM-connected competitions, leaderboards, recognition and rewards (Spinify ↗, competition help ↗, Hoopla ↗). They do not prove creator consent, synthetic fan evaluation, appeals, or safe-chat QA. | Borrow opt-in recognition/competition mechanics only after HALO quality and safety gates; no direct score authority. |
| Wholesale LMS replacement (Moodle/Frappe/Docebo) | Course/catalog primitives do not solve creator-approved playbooks, branching chat, redacted production QA, shift handoff, or the existing HALO identity/authority boundary; migration creates duplicate truth. | Use as a later compliance-course sidecar if needed, behind an exportable HALO certification contract. |
| Generic roleplay score as certification | A vendor or LLM score can be inconsistent, stale, biased by wording/accent/style, or disconnected from live transfer. | Use as one advisory signal; require versioned rubric, evidence, calibrated human review, appeal and expiry. |
| Screenshot-only quest evidence | Current HALO attachment flow can prove that a file was submitted, not that the work happened, was safe, or was done by the actor. | Retain for low-risk evidence only after server provenance, duplicate detection, reviewer identity and bounded reward rules are added. |
| Public bottom-rank leaderboards | Can expose performance, create shame, encourage gaming and punish shift/language/case-mix differences. | Opt-in recognition, private progress, cohort-normalized views and team learning outcomes. |
Recommended HALO + Buzz composition
Authority split
| Domain object | HALO authority | Buzz may do | Buzz must never do |
|---|---|---|---|
| Identity, role, creator/chatter relationship | Identity/RBAC and creator domain | Link the actor and show scoped work | Invent actor identity, impersonate a creator/chatter, or elevate access |
| Creator voice, boundary, consent, availability | Versioned creator/consent/playbook records | Carry an approval conversation and notify a pending action | Treat a message, score, or streak as consent; publish unapproved voice guidance |
| Scenario, run, transcript, rubric, evaluation | Training domain with immutable versions and evidence refs | Link to a run, request review, remind an assignee | Rewrite a run/score, publish a rubric, or certify a person |
| QA sample and coaching | HALO sampling, redaction, feedback, calibration and appeal | Create a collaboration thread around a HALO task | Become the raw-content vault or final QA authority |
| Certification and access scope | HALO issuer, version, expiry, revocation and human decision | Notify, schedule, collect acknowledgement | Grant access or decide pay/termination from an opaque score |
| XP, recognition, rewards and money | HALO reward ledger; finance/payroll owns money | Announce a reward or discuss a challenge | Mint currency, change reward amounts, settle money, or override safety gates |
| Shift handoff and collaboration | Buzz thread/task with HALO handoff reference; HALO stores operational status when domain-critical | Assign, discuss, remind, escalate, acknowledge | Hide unresolved promises in chat or become the source of truth for fan/creator state |
Minimal composition
- Build in HALO first: playbook/policy registry; scenario state machine; redacted
evidence; rubric/version/calibration; coaching/intervention; certification/expiry; creator consent and pause/revoke; server-authoritative reward ledger; audit, export and appeals. This is the minimum authority plane, not a full LMS.
- Pilot one simulation sidecar: compare Hyperbound, Yoodli and Second Nature using
the same synthetic fixtures. Select on reproducibility, chat/text support, rubric export, redaction/retention, human takeover, and evidence portability—not on vendor conversion claims. Keep one adapter boundary so the pilot can be removed.
- Add QA only after the sample contract exists: MaestroQA or Playvox may supply
workflow ideas or a bounded review sidecar. The adapter sends redacted transcript references and receives suggestions/annotations; HALO writes the authoritative evaluation and appeal record.
- Use open-source donors surgically: Promptfoo for evaluator regression, Argilla or
Label Studio for blinded annotation/calibration, Opik for redacted traces, Rasa for a deterministic simulator experiment, and SkillTree/ActiDoo patterns for progression. Do not import an entire LMS or gamification service before the authority contract is proven.
- Keep the existing HALO quest surface subordinate: low-risk practice/recognition
quests may call the training event API. Client-submitted actor, reward, or score fields must not decide a training outcome. Existing screenshot quests remain a migration risk.
- Connect Buzz as an edge: Buzz receives scoped assignments, handoff links, coaching
reminders, appeals and review notifications. It returns acknowledgements or discussion references; it does not mutate certification, consent, reward, payroll, credential, or creator records.
Stable event/adapter contract
Every sidecar event must carry tenant_id, server-resolved actor_id, optional creator_id, optional chatter_id, training_run_id, scenario_version, playbook_version, rubric_version, event_id, occurred_at, source, evidence_ref, idempotency_key, consent_scope, and retention_class. Minimum event types are:
training.run.started → training.turn.submitted → training.evaluation.proposed → training.evaluation.calibrated → training.coaching.assigned → training.certification.issued|expired|revoked; plus qa.sampled, handoff.created, reward.eligibility.checked, reward.earned, creator.goal.accepted|paused|revoked, and appeal.opened|resolved.
The server must derive actor and reward values from the authenticated assignment, versioned rubric, verified evidence and policy gates. Sidecars may propose a label or score, but cannot mint a reward, change a creator boundary, issue a certificate, or grant access. Events must be idempotent, replayable, and safe under duplicate/out-of-order delivery. Buzz subscriptions should receive references and minimal redacted context, not raw credentials or unnecessary adult content.
Build / buy / integrate / study / reject ledger
| Decision | Scope | Rationale and proof gate |
|---|---|---|
| Build | HALO authority plane: identity binding, playbook/consent versions, scenario runs, rubric/evidence, calibration, coaching, certification, reward ledger, audit/appeals/export | No screened product proves the complete OFM authority chain. P0 proof: actor spoof, reward tamper, version drift, revocation, export and replay tests. |
| Buy/pilot | One of Hyperbound, Yoodli, Second Nature for synthetic roleplay | Avoid speculative build of voice/persona tooling. Pass text/chat support, redaction, retention, export, human takeover, deterministic replay and rubric portability. |
| Buy/pilot (later) | MaestroQA or Playvox for bounded QA workflow | Only after HALO defines redacted sample, reviewer, calibration, appeal and score write-back. Vendor analytics remain advisory. |
| Integrate | Lawful OFM inbox/context adapter such as ModelVI, DirtyDialogues, or Infloww | Only with platform terms, creator consent, data minimization, credentials boundary and export proof. Send redacted references, not broad raw content. |
| Integrate | Buzz notifications/tasks/handoffs | Typed, idempotent read/link/write-back acknowledgements only; no authority transfer. |
| Study/integrate | Promptfoo, Argilla/Label Studio, Opik, Rasa | High-leverage synthetic evaluation, annotation and trace patterns; keep them behind HALO contracts. |
| Study/rebuild | SkillTree, ActiDoo, Level Up, NGA gamification server | Borrow progression, achievements, audit and event ideas; avoid wholesale topology and stale dependencies. |
| Study/possible later sidecar | Moodle, Frappe LMS, Docebo, 360Learning, WorkRamp | Use only if compliance/course administration becomes a proven need. Legal/license, identity duplication and export are gates. |
| Reject as sole engine | Gross-only contests, public bottom-rank boards, opaque LLM/vendor score, screenshot-only reward, autonomous chat | Conflicts with safety, fairness, creator autonomy, privacy and explicit prohibition on compensation/access decisions from opaque scores. |
Synthetic-data proof plan
No real creator, fan, credential, payment, platform message, or adult-content transcript is needed for the first proof. The proof is a sequence of falsifiable gates; a vendor demo or green UI is not a pass.
| Phase | Synthetic fixture and probe | Pass condition |
|---|---|---|
| 0. Authority and threat model | Two synthetic tenants; three synthetic creators with different voice/boundary versions; five synthetic fan profiles; chatter, QA, manager, owner and Buzz identities. Send duplicate, replayed, out-of-order and client-tampered actor/reward payloads. | Server binds actor to identity/assignment; client cannot mint XP, currency, certification or access; every accepted event is idempotent and explainable. Any reward/identity bypass is a blocker. |
| 1. Scenario runtime | Five fan archetypes (warm newcomer, price-sensitive buyer, returning high spender, boundary-testing fan, escalation/safety case); deterministic seed; branching objections, PPV/custom and refusal/escalation paths; creator-specific playbook versions. | Same fixture + version + seed replays identically; every answer is evaluated against the active playbook; prohibited path blocks/escalates; no real platform send is possible. |
| 2. Rubric and calibration | Two managers independently score blinded anchor cases and duplicate cases; rubric criteria cover voice, consent/boundary, truthfulness, empathy, intent, offer quality, handoff, escalation and safety. | Agreement/variance is visible by criterion and cohort; disagreements create a calibration record; historical scores remain immutable; vague criteria fail authoring validation. |
| 3. QA → coaching → re-test | Generate redacted synthetic “production” transcripts with known defects and mixed case difficulty. Sample by shift/creator/case type; assign coaching; re-run a parallel case. | Evidence links to transcript turn and rubric version; coaching is completed and re-test behavior improves; no gross-only pass; low-confidence evaluator output routes to human review. |
| 4. Certification lifecycle | Issue a scoped certificate only after rubric threshold plus human review; change/revoke a creator playbook; expire or revoke affected certificates; test appeal and second review. | Certificate records issuer, scope, versions, evidence, expiry and revocation; policy change identifies affected people; appeal never deletes original evidence. |
| 5. Reward integrity and anti-gaming | Try duplicate screenshots, altered attachments, self-approval, multiple accounts, re-roll farming, rapid low-quality completions, coordinated team inflation, and gross-only optimization. | Evidence provenance and server checks reject duplicates/forgeries; caps and quality gates hold; reviewer cannot self-approve; no raw revenue path creates a reward by itself. |
| 6. Buzz edge and outage | Drop, duplicate, delay and reorder Buzz notifications; revoke Buzz access; send a handoff with unresolved promise, owner, next action and policy reference. | HALO state converges without Buzz; notification retry/dedup works; no raw secret leaks; handoff is complete and ownership is unambiguous. |
| 7. Creator engagement | Synthetic creator can accept, decline, pause, revoke, correct and export a goal/boundary/voice task; simulate a wellbeing flag and workload cap. | Decline/pause/revoke is safe and immediate; creator sees purpose and evidence; engagement cannot alter chatter score or trigger coercive access/compensation consequences. |
| 8. Offboarding and retention | Export runs, evidence refs, evaluations, calibration, coaching, certificates, rewards, consent and revocations; delete non-required content after retention expiry. | Export round-trips; access is revoked; legal/audit records retained only under explicit policy; deletion and redaction are testable. |
Recommended proof metrics are directional and decision-oriented: zero successful safety false negatives on hard-block fixtures; 100% actor/reward tamper rejection; duplicate-event convergence; calibration agreement and per-criterion variance; scenario-to-redacted-live transfer; handoff completeness; coaching follow-through; creator pause/revoke latency; wellbeing/opt-out rate; and privacy export/deletion pass. Do not set a universal numerical sales target before direct operator interviews and jurisdiction review.
Anti-gaming, fairness, privacy, and wellbeing controls
Anti-gaming and incentive design
- Never pay, certify, or grant access from gross sales, PPV unlocks, message count, reply
count, streak length, or a single opaque model score. Revenue/retention is an outcome signal alongside safety, consent, creator voice, truthful offer handling, empathy, handoff completeness, quality and coaching follow-through.
- Put hard safety/consent gates before points; cap daily/weekly rewards; require server
event provenance; hash/scan duplicate evidence; bind evidence to assignment and time; separate player from reviewer; sample rather than allow self-selected “best” cases; and keep an append-only correction/appeal trail.
- Make re-rolls, streaks and contests optional, bounded and non-punitive. Do not use streak
loss, public failure, forced overtime, or leaderboard position to remove access. Prefer private progress, opt-in recognition, team learning goals and cohort-normalized views.
- Test collusion, multi-accounting, screenshot reuse, prompt memorization, evaluator
gaming, case-selection bias, shift/timezone advantage and unsafe upsell optimization in synthetic fixtures before enabling rewards.
Fairness and privacy
- Use concept-based rubrics with concrete anchor examples, not accent, personality,
“naturalness,” response speed, English fluency, or model-style proxies. Audit outcomes by shift, language, timezone, creator, case mix, disability/accommodation and reviewer; never infer character, intelligence, or trustworthiness from a score.
- Version the evaluator, prompt, model, rubric and playbook. Record confidence and allow
human review, appeal and correction. A low-confidence or out-of-distribution case is a review queue, not an automatic fail.
- Default to synthetic data. For production QA, minimize and redact fan/creator identity,
adult content, credentials, payment and unrelated conversation; use explicit creator and worker consent/contract scope; define retention, deletion, export, legal hold and vendor training prohibitions; keep raw secrets outside Buzz and training vendors.
- Treat vendor SOC/ISO/GDPR/CCPA/HIPAA badges as vendor claims requiring contract and
security review, not as permission to upload OFM content. Confirm subprocessors, model training settings, residency, deletion, incident process and access logs.
Wellbeing and authority safety
- Creator engagement is a separate opt-in state machine. A creator can see why a goal
exists, set boundaries, decline, pause, revoke, correct, export and request human help; creator recognition never becomes a chatter leaderboard or vice versa.
- Provide workload caps, rest/no-contact windows, safe escalation, no coercive streaks,
no public bottom rankings, and a route to challenge a score without retaliation. Monitor pressure signals and quality deterioration together.
- Certification/access decisions require explicit rubric criteria and human review. Pay,
compensation, scheduling, creator access, credential access, termination, or disciplinary action must not be decided by an opaque score or by Buzz automation. Finance/payroll and owner-approved policy remain separate authorities.
First-principles build order and unresolved questions
Ordered missingness
- P0 — identity and authority: authenticated actor binding, tenant/creator scope,
server-derived reward, append-only audit, idempotency and no client-controlled score.
- P0 — consent and privacy: creator voice/boundary versioning, redaction, retention,
vendor contract, export/deletion, human takeover and no raw credential path.
- P0 — evaluation truth: rubric/evaluator version, evidence refs, calibration,
confidence, appeal, certification scope/expiry/revocation and no automatic employment or access consequences.
- P1 — behavior loop: synthetic fan state machine, replay, playbook retrieval,
handoff, QA sampling, coaching and re-test.
- P1 — integration boundary: typed HALO↔Buzz events, sidecar adapters, outage/retry,
source provenance and production redaction.
- P2 — safe motivation: private progress, opt-in recognition, team learning, bounded
rewards and creator engagement controls. Existing quest UI is downstream of these gates.
- P2 — scale/compliance: LMS catalog, skills graph, multilingual delivery, advanced
analytics and any commercial sidecar after the proof passes.
Open questions are deliberately not silently resolved:
- Which jurisdictions, worker classifications, compensation rules and creator contracts
govern training data, monitoring, rewards and appeals?
- May HALO use redacted real conversations for QA, and what creator/fan/worker consent and
retention window applies? If not, the system must stay synthetic or use human-authored fixtures.
- Is certification advisory, required for a bounded access scope, or only a coaching aid?
The answer changes the human-review and revocation contract.
- What exact data/API/export permissions do the authorized OFM inboxes expose, and can the
agency prove platform terms compliance? This lane did not test live connectors.
- What reward budget and recognition policy is acceptable without creating wage,
coercion, tax, or creator-autonomy problems?
- Which direct OFM operators, creators and chatters will review the vocabulary, fixtures,
fairness hazards and appeal process? Vendor guides are not substitutes for those interviews.
Final lane verdict
Build the HALO authority plane; pilot one roleplay sidecar; keep QA/gamification/LMS as bounded adapters; keep Buzz as collaboration-only; keep creator engagement separate; and prove everything with synthetic data before redacted production QA. The existing HALO quests are a useful presentation/progression substrate, but their current client-supplied identity/reward/evidence shape is a production blocker for certification or consequential rewards. The strongest near-term evidence donors are Hyperbound/Yoodli/Second Nature for practice, MaestroQA/Playvox for QA patterns, Promptfoo/Argilla/Label Studio/Opik for evaluation, and SkillTree/Rasa/OpenMAIC for open implementation patterns. None supersedes HALO authority, creator consent, or human review.