Berlin startup Ojin launched on Product Hunt in late August, promising something the conversational AI market has struggled to deliver: photorealistic digital humans that respond to speech in real time. The company's pitch centers on latency below 200 milliseconds, fast enough that an AI agent can interrupt a user mid-sentence or react to tone before the person finishes talking.
Most avatar tools rely on pre-recorded video clips or voice-only pipelines. Ojin instead synthesizes facial expressions frame by frame, generating movement synced to live audio streams. The difference matters in practice. Slower systems create perceptible gaps when an agent reacts, the sort of delay that breaks immersion and reminds users they're talking to software. Ojin's approach tries to close that window.
The company offers two face models through a WebSocket API. Oris Portrait prioritizes speed and scale, hitting that sub-200ms target. A second model, Oris Presence, launched in July 2026, layers in what Ojin calls "maximum expressiveness"— micro-expressions and gesture fidelity meant to sidestep the uncanny valley. Both models start with a single still photograph and a live audio feed. Developers supply their own speech-to-text, language model, and text-to-speech components; Ojin orchestrates the timing to keep the face locked to the voice.
"Getting face, voice and timing to land in the same frame is a genuinely nasty pipeline problem," founder Christian Loclair wrote on Reddit earlier this year. "We built the chain ourselves instead of stitching together three suppliers."
That vertical integration reflects a broader bet. While competitors like HeyGen's LiveAvatar, D-ID Agents, and Soul Machines tackle similar challenges, Ojin positions itself as framework-agnostic infrastructure. The startup published integrations for Pipecat and LiveKit Agents, treating its API as a module that slots into existing voice-agent architectures without replacing the reasoning layer. You bring your own large language model. Ojin stays modular by design, the company has said in technical documentation.
The platform supports multilingual text-to-speech with voice cloning across more than 70 languages, Loclair noted on Product Hunt. He also cited company-claimed metrics including a 98 Net Promoter Score and average interaction times around 21 minutes, though these figures lack external benchmarking.

Loclair founded Journee Technologies, the legal entity behind Ojin, in September 2020, according to German commercial registry filings. He studied at Hasso Plattner Institute and the University of Potsdam. Before Journee, he founded Waltz Binaire and exhibited work at Centre Pompidou, ZKM, and Ars Electronica. The company rebranded to Ojin on April 13, 2026, signaling a pivot from broader metaverse ambitions toward real-time generative AI infrastructure.
Pricing starts at five cents per minute. Plans range from a Creator tier at $8 monthly (billed annually) to Enterprise, which includes managed deployments and dedicated support. New users receive $10 in signup credits. LinkedIn listed the team at 11 to 50 employees as of August 2026. Crunchbase indicates Journee Technologies raised over €20 million, though investor names and specific round details are not public.
The real-time conversational avatar market is splitting into distinct categories. Asynchronous video generators like Synthesia, which raised $200 million at a $4 billion valuation early this year, serve a different use case than live-interaction tools. Ojin competes in the latter space, where low latency and responsive expression matter more than polished pre-production.

The company claims compliance with SOC 2 Type II, GDPR, the EU AI Act, and PDPL standards. EU AI Act transparency obligations for AI-generated content took effect in August, adding regulatory pressure to an already complex technical challenge. Ojin runs demos from its AixHaus space in Berlin and maintains live web examples, including a Shopify product expert called Chiara and a careers-page Q&A bot. Named client testimonials on the homepage cite executives from H&M, Clinique, and BMW, though the scope of those deployments isn't detailed.
Whether the technology can scale beyond demos and pilot projects remains the open question for Ojin and its competitors. Fast facial synthesis solves one problem, but conversational AI still stumbles on context, reasoning, and the messy unpredictability of human interaction. A realistic face that reacts quickly doesn't fix a chatbot that misunderstands the question.
