I wanted to simulate characters reacting to what they hear. Each character would have its own profile, emotional state and way of showing a reaction. The idea is to feed it the words from a speech-to-text model. Generating speech can come later. First I wanted the listening side to work.
What does the character profile actually change?
If I describe someone as sensitive, reserved and concerned about losing other people’s respect, that should affect what happens when their work is criticised. How much does it affect them? Do they become angry, ashamed, sad? Do they show it? And what state are they in when the next sentence arrives?
That is what I am building with Character Reactor.
Take two fictional characters from the prototype. Mara is a meticulous stage designer. She cares strongly about dignity and belonging, has low assertiveness and tends to keep her reactions contained. Leo is a confident improviser. He values autonomy, is more assertive and recovers more quickly after tension.
Now give both of them the same sentence:
Your work is careless and disappointing.

I tested this sentence with both profiles. The model classified it as criticism for both characters. Mara’s simulated state rose most strongly in sadness, with shame and anger also increasing. Her chosen response was to withdraw. Leo’s strongest increase was in anger, and his response was to hold his ground. He showed more of that reaction.
These differences are partly designed into the profiles. Mara’s description already says she tends to withdraw; Leo’s says he holds his ground. The numeric traits and needs also feed the state calculations. This run shows those authored profiles producing different consequences. It doesn’t show a model discovering two personalities on its own.
I followed it with an apology:
I’m sorry. I was unfair to you.
The system recognised an attempt to repair the interaction. Negative emotions eased as time passed, with a small additional reduction from the repair signal. Trust increased slightly. Neither character returned to its starting state.
I placed the apology seven seconds after the criticism, then repeated the criticism another seven seconds later in the simulation. Both characters ended with more negative emotion than the same sentence produced in a fresh session. Mara withdrew again. Leo held his ground again. What changed was the emotional state behind those actions and the intensity of the visible response.
That is what I wanted to make possible: a character whose next reaction depends on what has already happened. Here it happens through accumulated emotion, changing trust and different recovery rates. The prototype does not yet recognise a pattern such as “you apologised and then did it again.” That would require the remembered interaction to become part of the next judgment.
What the profile changes today
I started with sensitivity, assertiveness, composure, expressivity and optimism, alongside the importance of safety, dignity, belonging and autonomy. The profile also gives the character a starting emotional state, a recovery rate and trust values for particular speakers. Description and goals help select how it responds.
Criticism affecting dignity raises anger more strongly for an assertive character and shame more strongly for one with low assertiveness. Lower trust in the speaker increases fear from danger or threatened belonging. Composure and expressivity affect how much of the emotional state becomes visible. Recovery determines how much of an earlier event remains when the next one arrives.
These are rules I have chosen for a simulation. They give me something explicit to inspect and change. They are not a validated account of human psychology.
The current limitation is that both characters receive the same initial interpretation of the sentence. The model reads the words without their individual beliefs or remembered experiences. Character differences enter through the state calculations and the later choice of response.
For the fuller idea, the profile needs to influence interpretation too. A correction from a trusted colleague could mean something different from the same correction delivered by someone who repeatedly humiliates you. A character trying to protect a collaboration might hear a risk to the relationship where another hears a challenge to its independence. Those are the distinctions I want to model next. Beliefs and descriptions of relationships are already written into the profiles, but the current decisions do not use them.
How the character reaches a response
I wanted to do this with Type 1 models. By that I mean models that answer bounded questions, choose among defined options or return a score. The system produces structured reactions. It does not generate dialogue or speech.
I split the reaction into three steps.
First, the model identifies the kind of event: criticism, reassurance, a request, a threat, or something else. That choice determines which questions to ask next.
Second, it assesses the relevant dimensions of the message, such as danger, rejection or support. The simulation code uses those judgments, the profile and the current state to calculate emotional changes.
Third, the model chooses a way of responding and how much to show, using a short description of the character, its goals, a rough description of its current feelings and the event category.
The response can be withdrawal, assertion, clarification or another permitted action. The chosen action and emotional state are expressed through posture, gaze and facial parameters.
Each model call produces a defined choice, a score against a rubric, or support for a yes/no statement. For the browser version, I adapted a DeBERTa text classifier to those tasks. It scores how well the input supports a set of statements we have supplied, such as “This text contains an insult or criticism.” The questions and available responses are part of the design.
The hierarchy is useful because each step has a job I can examine. If criticism produces a weak reaction, I can check whether the message was misclassified, the emotional impact was too small, or the face failed to show the state. Adding another level does not by itself make the character more convincing.
I ran into exactly that problem with an earlier version. A direct insult hardly moved the characters, and the faces were too inexpressive. Improving it involved changes to the semantic model, the emotional dynamics and the facial mapping. All three mattered to the design; I have not measured their contributions separately.
The face needs to show more than one emotion at a time. Anger can sit alongside shame or sadness. A character can be affected strongly and still show very little. The code combines the model’s display score with the emotional state, composure, expressivity and the chosen response. It maps those into 28 continuous facial and pose channels, including brows, eyelids, mouth shape, gaze and head position. Those mappings are authored, so they can be inspected and adjusted. A stronger display makes the state easier to read; it does not establish that the emotion itself is correct.
Streaming input adds another practical problem. A transcript can change while a person is speaking. “I hate…” might become “I hate what they did to you.” Applying every partial version as a lasting event would leave the character reacting to things the speaker never finished saying. Character Reactor previews a reaction while a segment is still changing, then updates lasting state once the segment is final. Repeating that final segment does not apply the emotional impact again.
The live microphone uses Vosk’s small English speech-recognition model through WebAssembly in the browser. It previews partial transcripts and commits each segment when the recognizer marks it final. The app, runtime and models total about 180 MB. Use “Save for offline” to download and cache everything needed for an offline reload. The app does not generate a voice. I checked this path in Chromium with a synthetic English speech fixture; that verifies this test setup, not general recognition accuracy or performance across browsers and devices.
I also wanted to make this easy to put on a web page. The browser package uses an open-source English DeBERTa model, with quantised weights and CPU inference through WebAssembly, using Transformers.js. There is no inference service to operate; the visitor’s device handles the model download, memory use and computation.
The recorded sequence used that WASM inference provider under Node. It demonstrates this implementation running through those three events; it is not a browser performance measurement or evidence that the characters behave like real people. The scores still need calibration, and testing longer interactions is part of the work ahead.
Mara’s profile says: “An apology matters when behavior changes.” Today that belief has no effect on the model’s judgment. It describes the character, but the system does not use it when deciding what the repeated criticism means.
That is what I want to test next. The words stay the same, but the character has heard an apology between the two criticisms. Its belief about apologies should affect how it interprets what happens next. I would compare the judgment with and without that belief, keeping the other profile values and the starting state fixed. I want to be able to follow the difference back to the character’s profile and experience, and see whether it makes sense.
Try Character Reactor · Download the MIT source · Explore both experiments