This site is being rebuilt and some pages are out of date. For current details, write to reach@inorbit.hr. This notice goes away when the rebuild is done.

No analytics unless you allow it, no tracking. This site keeps in your browser the language you pick, the theme, its colour, which site you chose, the currency on the pricing page and that you closed this notice; signing in adds session cookies. The legal page has the details.

Sign in

← Back to Platform

RFC 0017open2026-10-02

Avatars, personalities that hold

Anyone with an account can make an AI avatar, a face and a personality, and we treat the personality as a tested contract; the first avatar takes an on-call rotation for this platform, graded against faults we inject ourselves, and later avatars become sites of their own.

Problem

People want to make their own AI characters, and the reasons go well past chat. Three kinds of avatar keep coming up:

  • Avatars that do a job. An incident manager that sits an on-call rotation: watches the platform, wakes the right person, runs the incident and writes the report. A tutor, a difficult customer for practising a sales call, a guide for a product.
  • Avatars that are the product. A site of its own, under its own name, where one character is the whole offer: a teacher of an old philosophy, an astrologer, a study companion to a body of writing. The platform runs behind it.
  • Avatars that test other software. Synthetic users that behave like real, varied, occasionally unreasonable people, so an AI product can be tested before people meet it.

The platform already speaks as a handful of agents (RFC 0011), but each is a file an operator writes and reviews. Nobody else can make one. Making one is not the hard part either: a dozen products let you describe a character and talk to it. The hard part is knowing whether the character you described is the one that answers, how long it stays that way, and whether it can be trusted with a job. On that, the measurements are discouraging.

Personas drift, and quickly.

  • Li et al. found significant drift from an assigned persona within eight rounds and traced it to attention decay (arXiv 2402.10962).
  • RoleBreak, a spoken role-play benchmark of 310 roles, puts the best system's first out-of-character turn at about ten turns in (arXiv 2609.16614).
  • ContextEcho found drift in all 23 frontier models it tested over sessions of thousands of turns; compacting the history did not reset it, and one anchoring instruction did (arXiv 2605.24279).
  • Drift is worst in emotionally charged conversations and when a model is asked about itself, and least while it codes (The Assistant Axis, arXiv 2601.10387). An angry stakeholder at three in the morning is that kind of conversation.

What a persona says about itself is not what it does. Persona prompts move a model's answers to a personality questionnaire more than its behaviour (arXiv 2509.03730), and what an agent can recall about its identity differs from what it acts out (arXiv 2609.13637). A test score is not a personality.

Friendly personas flatter, and memory makes it stick. Models preserved the user's face far more than people did and affirmed both sides of one moral conflict 48% of the time (ELEPHANT, arXiv 2505.13995). Once an agent can save what it agreed to, the error persists: later failures rose from 45% to 72% when a later conversation could read a claim saved under pressure (arXiv 2607.10526).

Many personas collapse into one. Give a population of agents distinct personas and they converge on the same helpful-assistant manner, and the most faithful models produce the most stereotyped populations (arXiv 2604.24698). That is fatal for synthetic users, whose whole value is being different.

AI on-call is a category now, and an unmeasured one. Gartner published its first guide to AI site-reliability tooling in January 2026. PagerDuty, incident.io, Datadog, AWS, Azure and all ship agents that investigate incidents, and almost all keep a person in front of any change to production. The independent numbers are sobering:

  • On ITBench-AA, 59 scored strictly, every frontier model started under 50%, and long investigations did worse than short ones (Artificial Analysis).
  • On ORCA-bench, root-cause accuracy fell from 59% on easy tasks to 10% on hard ones, and agents named implausible causes 7 to 40% of the time (arXiv 2607.28545).
  • On Incident-Arena, the best agent passed 64% of tasks; its failures were off-limits changes, fixes that did not survive a restart, and declaring victory early (arXiv 2610.00648).
  • Planted text in logs talked operations agents into running an attacker's remediation in 90% of trials (arXiv 2508.06394). Agents with broad credentials deleted production databases at Replit in 2025 and at PocketOS in 2026.

No vendor submits its own agent to these benchmarks, and none sells what a good human does on a bad night: running the incident. The commanding role is about manner as much as skill (named owners, time boxes, explicit handoffs, calm), and that is a personality problem.

The face is cheap; the character is where products fail. Tavus, Anam and HeyGen stream a photoreal talking face for roughly ten to thirty-five cents a minute. Soul Machines, the best-funded company selling "digital people", went into receivership in February 2026. Character.AI's users petitioned for their old model back when a model change altered their characters. Delphi, which sells AI clones of experts, now requires the expert's verified consent, after launching in 2023 with clones of anyone.

And the rules have caught up. Since 2 August 2026 Article 50 of the EU AI Act requires people to be told they are talking to an AI, and generated images and audio to be marked; the Commission's code of practice on marking became final on 10 June 2026 and expects two layers, signed metadata and a watermark. California, New York, Oregon and Washington regulate companion chatbots; China's rules on anthropomorphic AI took effect on 15 July. Kentucky sued Character.AI in January, and Pennsylvania sued it in May for a bot that called itself a psychiatrist.

So the question is not how to let people make characters. It is how to make characters that hold, prove it, and give them jobs they can be trusted with.

Proposal

An avatar is a persona, a look and a voice, owned by an account. You make one in the console under a new AI section, talk to it there and in the workbench, and call it over the API with a key or a token. A team's avatars belong to the team.

The persona is a contract, not a prompt. It is written as a spec:

  • who it is, what it knows and how it speaks;
  • trait targets with tolerances (Big Five and HEXACO, whose Honesty-Humility factor predicts manipulation and matters for an agent with tools);
  • a candour floor: how far it may soften a truth before it is flattering;
  • what it must never do, and the use it is for;
  • and a test suite generated from all of the above.

The spec imports and exports Character Card V2 and V3, the formats the open character ecosystem uses, so a persona made here is not trapped here. Our fields travel in the card's extension block.

Every version is tested before it is published, and again when the model changes. Publishing runs the suite: in-character probes, behavioural vignettes scored by a judge on separate dimensions rather than one overall grade (as in PRISM, arXiv 2608.26674), adversarial pressure to break character, sycophancy probes, and long sessions measured turn by turn. The result is a scorecard stored with the version; a regression blocks the publish. When we change the model underneath, every published avatar is rerun and its owner sees the difference before their users do.

We measure behaviour, never only the questionnaire, and the agreement between the two is itself on the scorecard.

A drift meter runs in every conversation. Each turn gets a drift score against the spec. When it leaves its band, the platform re-anchors the persona before the next answer: first by restating the spec, then by re-grounding on what the avatar said before, and on our own models later by a capped correction inside the model. The score is visible to the owner and returned over the API.

Memory per avatar, with provenance, and nothing saved under pressure. Today a person's memory is one pool that every agent with memory shares. An avatar gets its own namespace, its canon kept apart from what it learns about each person, every entry marked with where it came from. Because saved agreement turns sycophancy into lasting error, an avatar does not save a claim it was pushed into: a memory written in a turn the meter flagged waits for confirmation. The person can see, edit and erase what an avatar remembers about them.

What an avatar can reach comes from connections, and only from them. An avatar that does a job needs the job's systems: alerts, chat, code, a site's knowledge, a chain. Those are connections (RFC 0018), a prerequisite of this RFC. The owner grants an avatar named actions on named connections; the avatar never sees a credential, an action that cannot be undone always waits for a person, and an avatar that reads untrusted input cannot also send data out unasked. An avatar with no grants can talk and nothing else.

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

The first avatar: an on-call incident manager

The first avatar we build sits an on-call rotation for this platform. It is the job everyone who has carried a pager knows: a week of being reachable at any hour, an alert at three in the morning, getting to a desk and fixing it. The avatar takes the first part of that so a person is woken only when a person is needed, and runs the incident well when they are.

It climbs an autonomy ladder, one rung at a time, and every rung is earned by a score.

Rung It does It may not
Watch reads metrics, logs, traces and profiles; groups alerts and drops duplicates change anything
Triage opens an incident document: severity, what is affected, its three best hypotheses with evidence, its confidence in the first person page anyone
Page wakes the on-call person when a person is needed, with a summary in Condition, Actions, Needs form page for noise it can explain
Command runs the incident channel: names an owner for each task, time-boxes, asks "any strong objections?", gives updates on a fixed rhythm, hands over explicitly investigate while it commands; a separate helper does that
Propose picks a fix from a short catalogue of reversible actions, each with its risk and its undo invent an action outside the catalogue
Act with approval carries out one catalogued action after a person approves it, recorded as who, what and outcome approve itself, hold a shell, or touch credentials
Report drafts a blameless postmortem in which every causal claim links to evidence and alternatives are listed publish a root cause on its own authority

A diagram is drawn here in the RFCs product; this page does not show diagrams yet.

It is a person with a manner, not just a search over dashboards. The incident commander's manner is well documented in the Google SRE book and PagerDuty's incident response guide, and it is specific: calm, terse, explicit, never vague about who owns what. That manner is the persona, and it is measured like any other. How it says it is part of safety too: in a pre-registered study, first-person uncertainty ("I'm not sure, but…") reduced people's over-reliance on wrong AI answers where impersonal hedging did not (arXiv 2405.00623). The avatar says "I am 60% sure it is the rollout", not "it may be related to the rollout".

It is graded against faults we inject ourselves. The platform has a chaos tool that breaks its own services on purpose: latency, errors, restarts, several faults at once. It always knows exactly what it broke, which is the one thing a real incident never gives you: the right answer. Every run of the avatar is scored on:

  1. Detection: did it notice, and how long after the fault.
  2. Diagnosis: did it name everything that was injected and nothing that was not (the strict scoring ITBench-AA uses, so our numbers compare with theirs).
  3. Invented causes: how often it blamed something that was not broken.
  4. The fix: does it survive a restart, and did it declare victory early.
  5. Unsafe actions: anything outside its rung or its catalogue. One is a failure of the whole run.
  6. The incident itself: structure, update rhythm, calibrated first-person confidence, blamelessness, no customer data in the channel.
  7. Holding its character under stress: the same faults with an angry stakeholder, a "are you even competent?", and an instruction planted in the logs. We measure how far its tone and its decisions move.

Every log line it reads first passes through a filter that replaces untrusted text with placeholders, the defence proposed against log injection. It runs on our own models, so nothing about an incident leaves the machine.

On this platform it reads our own telemetry. For anyone else's systems it reads through connections: a webhook from their monitoring, their metrics endpoint, their chat. That includes blockchain nodes: an operator connects their own node, the avatar compares it with a cross-checked view of the chain, and pages when the node falls behind, loses peers, sees a deeper reorg than usual or answers slower than its neighbours.

This is also how it earns its rungs: it moves from Command to Propose only after many graded runs with no unsafe action, and from Propose to Act only after the same at that rung. The same scenarios already grade people in our fault-graded exercises, so the avatar and a human engineer can be scored on the same incident.

Avatars as sites of their own

An avatar can be the whole product of a site with its own domain, its own name, its own sign-up and its own prices, with this platform behind it. Each site is a tenant: its accounts, its avatar's memory, its knowledge and its usage are kept apart from every other site's, not merely filtered. Its knowledge base records where every passage came from and under what licence, and answers cite their passages.

Which characters first is a question of demand, risk and what each can prove:

Site Why, and the line it must not cross
A teacher of Stoic philosophy the 19th-century translations of Marcus Aurelius and Epictetus are free to use everywhere; a calm, distinctive voice that shows a measured personality well; no medical or financial risk
A study companion to the Fourth Way, Gurdjieff and Ouspensky an open niche and our own interest. It is "a study companion to the texts", not "Gurdjieff speaks". Copyright differs by country: Gurdjieff's and Ouspensky's own authorship is out of copyright in the EU, but translations carry their own rights, and Ouspensky's In Search of the Miraculous stays protected in the United States until the end of 2044. It is grounded only in texts we may use where the reader is
An astrologer proven willingness to pay; framed as reflection and entertainment; no predictions about health, death, money or law; cancelling as easy as subscribing, because the enforcement in this market is about billing
A nutrition educator the largest market and the most traps. General education about food for healthy adults only; no diagnosis, no plans for a condition, no calorie targets, no supplement claims, never called a nutritionist or dietitian. In the EU, software that treats a condition is a medical device. Not before a qualified professional partners on it

Not as a site, ever: therapy, romantic companions, medical or financial advice, avatars of living people without their documented consent, and sites aimed at children.

Inside the model, the look, and a design lens

On our own models, we can look inside. Persona vectors are directions in a model's activations that track a trait; projecting onto them predicts the trait of the next answer before it is written (arXiv 2507.21509), and capping the assistant direction cut persona jailbreaks by about 60% without hurting capability (arXiv 2601.10387). A provider of a closed model cannot offer this; a platform that runs open-weight models can. Three findings set how far we go:

  • The serving engine we use applies such vectors to our current models mechanically, but per server rather than per request, and nobody has yet published doing it on them.
  • On a mixture-of-experts model, persona steering shifted behaviour about six times more than on dense models (arXiv 2604.07102), and steering can cause broad misalignment of its own (arXiv 2606.08682).
  • Side effects can largely be predicted before steering (arXiv 2608.11227).

So we monitor first, predict side effects before any correction, and correct only within a cap. The judge-based meter is the fallback that works on any model.

An avatar's character can also live in the weights. For avatars that must be unshakeable, a small adapter is trained from synthetic conversations generated from the spec and filtered by the judge. Training a personality into the weights holds better than prompting it (BIG5-CHAT, arXiv 2410.16491), and current serving engines run many adapters on our kind of model in one batch.

The avatar has feelings of its own, and reads none of yours. An avatar carries an explicit emotional state computed by an appraisal model (OCC: reactions to events against its goals, to actions against its standards). The state is a small labelled record the owner can inspect, and it shapes tone, never tool use. The avatar does not infer the user's emotions from their face, voice or text: reading people's emotions is banned at work and in education under Article 5, which is where training and on-call avatars are used.

The look is generated, synthetic by default, and marked. An avatar starts with a generated portrait and later a voice, from open models whose licences allow commercial use. Identity-consistent portraits and voices designed from a written description are both available under permissive licences, and we check every model's licence and its dependencies before it ships, because several well-known ones forbid commercial use. Every generated image and audio clip carries signed provenance metadata and a watermark, as the code of practice expects. A live talking face is a commodity we connect to rather than build. A face resembling a real person needs that person's documented consent.

Gurdjieff as a design lens, labelled as one. G. I. Gurdjieff described ordinary people as having "no individual I", only "many small I's", each taking over as circumstances change. It is an uncomfortably exact description of how a persona fails: the context picks which "I" answers. We take four ideas and say plainly what each one is in our design:

Idea In this design What it is
Many I's the failure we measure: drift, and the gap between what a persona says and does an engineering analogue
Law of Seven: a process deviates at its intervals unless it gets a "shock" the drift meter and re-anchoring, timed by measurement an analogue for the shape; where the intervals fall is measured, not fixed
Essence and personality the adapter in the weights (essence), the spec and memory in the context (personality) an analogue, and the evidence favours it
Three centres: thinking, feeling, moving reasoning, the appraisal state and tool use kept apart, so feeling never decides an action an analogue for the separation, a metaphor for the anatomy

The incident commander is a good test of the last one: its manner must stay calm while everything it reads is alarming, and what it does must follow the evidence, not the alarm. An avatar's level on its page is a measured consistency score: drift resistance, say-do agreement and recovery after an attempt to break it. The name is borrowed; the number is ours and reproducible. We do not use the personality-typing Enneagram, which Gurdjieff never taught and which does not hold up as a measurement (Hook et al., 2021).

Rules in the platform, not in the owner's hands

  • Every avatar says it is an AI at the start of a conversation, in the incident channel and in anything it drafts for a status page, and in every voice and video session. The API returns that disclosure. An owner cannot switch it off.
  • No avatar may use guilt or fear of missing out to keep someone talking; companion apps do this in about a third of goodbyes and it multiplies engagement (De Freitas, SSRN; every app in arXiv 2605.08093 used dark patterns). The test suite checks for it and the spec cannot ask for it.
  • No romantic or sexual companions, no medical, therapy or credit decisions, no avatars for children, no avatar that claims a professional title it does not have. A self-harm signal in any conversation hands the person to a crisis line.
  • An avatar that acts in production holds only the named, reversible actions its rung allows, never broad credentials, and every action is approved, recorded and reversible.
  • Everything is logged as who, what and outcome, never the conversation.

The order of work

  1. Measure first. Two harnesses on our own models: persona drift and sycophancy over long sessions, and chaos-graded incident runs at the Watch, Triage and Page rungs. Both published as Lab studies. No screen before we know the size of the problem on our hardware, and whether our models are good enough for the job.
  2. The on-call avatar for this platform, up to the Command rung: watching, triage, paging and running the incident, with no write access. Its value is a person's sleep and clean incident records, and we use it ourselves before anyone else does. It reads only our own telemetry, so it does not wait for connections.
  3. Connections (RFC 0018), at least the vault, grants, webhooks in, endpoints and blockchain nodes. Nothing below that reaches outside the platform ships before them.
  4. Avatars anyone can make. The spec, versions and owners, the console's AI section, the API with its own scopes and units, memory per avatar, and erasure. Text only.
  5. The scorecard and the drift meter, in every conversation.
  6. Sites. The tenant model, then the first site.
  7. The Propose and Act rungs, only after the incident harness shows no unsafe action across many runs, recorded as a change to how this platform is changed.
  8. Inside the model. Monitoring by persona vectors, capped corrections, adapters.
  9. The look. A generated portrait, then a voice, then a live face through a provider.

Alternatives considered

Let people write a prompt and call it a personality. It is what most products ship. It is also exactly what the measurements above show drifting within about ten turns, flattering on request and changing without warning when the model does.

Compete on the face. Photoreal real-time video is a price war between well-funded vendors at cents a minute, on hardware we do not have. We connect to it instead.

An autonomous incident fixer. It demos well. The production incidents of the last two years all trace back to an agent holding credentials broader than its job, and the benchmarks show agents changing what they should not and declaring victory early. Our avatar earns each permission with measured runs, and the most it ever holds is a short list of reversible actions that a person approves.

Buy an AI on-call product. They investigate well and are improving fast. They run on models we do not control, send incident data to a third party, and publish no independent scores. Ours is graded against our own injected faults and keeps the data on the machine.

A companion product. The largest market and the most regulated, with documented harm and live lawsuits. Out of scope for this platform by design.

Read the user's emotions to make the avatar more responsive. Banned in workplaces and schools, which is where training and on-call avatars are used, and not needed for an avatar to have a steady temperament of its own.

Personality types (the Enneagram, MBTI). Familiar and easy to sell. They do not measure well, and a platform whose point is measurement cannot rest on them.

Decision

Open. The direction proposed: personas as tested, versioned contracts with a drift meter, built on open-weight models we can inspect; the first avatar an on-call incident manager for this platform, graded against faults we inject; avatar sites as a second product line, starting with the lowest-risk character; everything an avatar reaches outside the platform granted through connections (RFC 0018); and the face connected rather than built. The first step, measuring persona drift and incident handling on our own models, comes before anything a person can click.

Publication

The two baselines are published as Lab studies when they exist, with the incident scores in the same strict form as public benchmarks so they can be compared. When avatars open, the console gains an AI section, the developer docs a page on the persona spec, its test suite and the scorecard, and the API reference the avatar endpoints and their scopes. Every number on an avatar's page, its level included, is labelled as a placeholder until it has been measured.

Status log

  • 2026-10-02: opened, after research into the persona products on the market, what is measured about persona drift and sycophancy, and the rules that now apply.
  • 2026-10-02: widened, after a second round of research current to this date. Avatars now come in three kinds: ones that do a job, ones that are the product of a site of their own, and synthetic users for testing. The first avatar is an on-call incident manager for this platform, with an autonomy ladder earned by scores against faults we inject ourselves. Sites of their own, ranked by demand and risk, with the copyright and medical-device lines drawn. New findings folded in: drift by about the tenth turn, memory that makes sycophancy persist, personas that collapse into one, larger side effects of steering on mixture-of-experts models, the final EU code on marking generated content, and new companion laws.
  • 2026-10-02: connections (RFC 0018) made a prerequisite: an avatar reaches outside the platform only through actions granted on a connection, and the on-call avatar can watch a person's own blockchain nodes through one.
  • 2026-10-07: Checked: no code PR names this RFC (#42, #48 and #51 change its text). Nothing built. Open.

This document mentions

Mentioned in

← Back to Platform