This Q&A has been prepared by the current development team.

We joined Maghni during its development process, and had no prior relationship with either Natalia Shmueli or Austin Grissom. We worked with both project leads and their respective companies, Misbah and VocaTone, during this time. Neither former lead directs or assists with our current development work.

This article covers events we personally observed and information we can corroborate through our records. Where ownership, contractual responsibility, or other legal issues remain disputed, we do not attempt to adjudicate them here.

We have attempted to cover as many questions as we could in this summary and Q&A. For future inquiries, we're setting up a comment section where questions can be placed.

Backstory & Timeline

In December of 2020, Natalia Shmueli and Austin Grissom announced Maghni AI, a project they jointly led, which had been in development for a year. This project was funded by Natalia. At the start of the project in December 2019, they posted an advertisement online to hire a programmer skilled in audio technology which resulted in Barham’s hiring.

Barham worked on concatenative synthesis during the first year of Maghni AI, eventually switching to AI synthesis soon after the December 2020 public reveal. As more technically-skilled people joined the team (Caleb and Tart in 2021, Aku in 2023), the team attempted to have Barham upload his code multiple times; a normal practice within software engineering, referred to as "source version control". After significant effort, we received only the old concatenative code, which was no longer relevant or usable.

By December of 2023, Barham worked on the AI engine far longer than anticipated. Natalia, who had provided the majority of the project's funds to that point, could no longer maintain Barham's previous payment schedule. Around this time, the team learned that Barham needed continued proof of financial support, as this was connected to his German Immigration status. This was presented to us as a serious threat to his ability to remain onboard the project.

In an attempt to save Maghni AI, a crowdfunder was launched in partnership with VocaTone, a company that was owned by Austin Grissom. It was agreed upon that the merchandise rewards for the crowdfund would be purchased utilizing the crowdfunding proceeds, with the majority of it going towards Barham in order to keep him paid and able to work. Progress proceeded steadily after this, though Barham often had month-long delays between updates. To give Barham additional support, Natalia and another developer bought an AI training machine (or "rig"), made to speed up the training of AI models. This rig was physically delivered to Barham in Germany by mail.

Natalia then left the project for personal reasons while remaining available to the team when contact with Barham needed to be facilitated.

Communication with Barham became increasingly sporadic, despite the team’s efforts to stay on task. July 9th, 2025 was the final time we received any sort of code from Barham. All ongoing training tasks were halted, along with our main developer and all of the code for Maghni which he had still not uploaded. We had received one additional piece of code since the original concatenative code; the inference script that let us use models Barham trained to output audio. We had no way to train our own models for Maghni. For reference, this is the first time the Maghni team had received any deliverables relating to the AI synthesis engine.

The AI engine was provided to us in an incomplete state where no additions or optimizations could be made. The team was required to redevelop the engine nearly entirely from scratch, working nonstop on code that had taken Barham over four years to develop, all while the team continued to balance personal responsibilities unrelated to the project. At this stage, even if the missing code were provided tomorrow, it wouldn’t be of much use given how much we've redeveloped the engine on our own, since the architectures remain largely incompatible.

At this time, we are currently working on rebuilding the entire AI engine and have come a significant way towards this goal. The current team remains dedicated to finishing the project. We are not associated with any of the prior team leads (Natalia and Austin), and will be rebranding the engine in the future to reflect this reality.

We hope this Q&A helps to answer some questions you may have. We will continue to answer questions on our personal socials, as well as the official rebranded software socials when they are created.

General Questions

There have been claims that renders created during the crowdfunding process have used complete engine output from other products. We believe these claims come from a general misunderstanding on how inference with the first version of the engine functioned. During the crowdfunding campaign, the editor front-end was unable to be given to producers or tuners for demo creation. The inference code remained on Barham’s machines, so integration with the editor was impossible. Instead, we first had artists give a project file with MIDI notes and the corresponding words. This could have been any MIDI editor; most chose to use Synthesizer V due to ease of access. After this, we calculated the length of each word and generated timing with our own timing system. The output was a MIDI file where each MIDI note represented a phoneme that would be rendered. Individual note lengths and phoneme fields using Maghni's systems were also included with the MIDI information to pass along to the inference pipeline. First-pass pitch curves were also provided, which developers would hand-copy into the engine. An end-to-end synthesis pipeline of this nature cannot accept audio as an input, only phoneme tokens, durations, and per-frame pitch. Artists and demo makers simply needed a solution for the initial MIDI creation process and phoneme duration editing. Images shared from the crowdfunding period were not accurate of the actual files we gathered information from. You can think of it more as a MIDI file, where each note is a phoneme, and the lyric data of the MIDI note is the phoneme to be rendered. This is also why demos were costly on time to make, since manually generating all the information to pass into the engine was hard without a fully linked editor.

For Oliver's "Right Where You Are" demo, RVC was first utilized to create a small speaking part. The engine output was unstable at this time, and under supervision from management this operation was greenlit. Two additional cuts of the render were made; one utilizing RVC-generated consonants and high-notes, and one with consonants from Oliver's voice provider's source files. In the final product, the RVC-edited render was used, with 60-80% of the audio remaining as raw Maghni output.

Please note that no other demonstration of Maghni was edited in this way.

Audine's voice was entirely provided by a private singer who was paid in full by Natalia. For her demo, the songwriter used an instrumental provided by a "free beats" YouTube channel; we did not know this nor was this disclosed to us. Upon learning this information, a disclaimer was attached to the video.

Despite our numerous requests for repository access and source code delivery, the team only received only two portions of the code; the inference code, for generating audio from models, and the very old concatenative code. While we are currently working to reconstruct a new engine inspired by the old implementation, the old concatenative code was not useful to the project.

It was sent to Barham physically for ease of training use. To our knowledge, the rig is currently in Germany. At no point in development after the purchase of the training rig, was it made available to any other team members.

Our engine’s primary software and hardware have been recreated multiple times. Despite our best efforts, we've rebuilt the code from near-scratch three times now; the current iteration of the engine is less than a year old. We wish we didn't have so many setbacks, but we are determined to keep going and finish this project.

Crescendia owns full rights to the recordings of their vocal libraries and engine products. Any ongoing production of their libraries is halted until further notice.

Crescendia did not receive or control the original crowdfunding proceeds, nor do we control VocaTone’s physical inventory, shipping process, accounts, or refund system. We cannot issue refunds for money we didn’t receive.

VocaTone and Misbah publicly presented and internally operated the Mahgni AI engine as a joint project. After VocaTone ceased their participation in development, our work continued under Misbah and the Crescendia group. Natalia eventually stepped back from active development, while remaining available for communication with the team when pertinent information was needed, or when contact with Barham was required. As a note: Our Q&A does not attempt to resolve legal ownership disputes between former parties.

At this time, our development team is going through the legal steps required to separate our work from both entities. We appreciate your patience as we finalize this process.

Our current team has no further association with the original project leads. We believe completely renaming Maghni would be the best way to showcase the change in ownership. We’re currently taking time to create entirely new branding.

We will keep everyone informed on our future changes. For sake of transparency, please understand that the engine as a product is entirely different from what it was during initial funding, and only retains libraries in name.

We're planning on launching a public Discord server within the next couple of months. This will include free, public access to news and updates. There have been many requests about wanting more direct contact with the current team, so we are aiming to provide frequent updates and communication. We are still in the process of making sure the server will be properly staffed and moderated and will announce more information at a later time.

The team remains passionate about the work they've done over the years and are confident they can release a useful product.

Current Engine Development Questions

  • Caleb, known by his screen name ccodaze or nickname Cocoa, is the director, backend developer, and general jack of all trades.
  • vocatart, known by his nickname Tart, is the lead library, scripting, and linguistics developer.
  • Miguel, known by his screen name Aku, is the lead design, branding, and frontend developer.

Along with other team members who prefer to remain unnamed.

We've only been working on this iteration of the engine for about two-thirds of a year. We cannot provide even a rough timeline, as we don't want to have to break a promise once again. We will show samples as soon as we can.

We received a very small portion of the acoustic articulation network (enough for simple inference). Not enough to integrate it into the editor, and not enough to train our own models. Rebuilding the other components would’ve been a misplaced effort, since R&D into a new stack would take the same amount of time while yielding considerably better results with newer technology at our disposal. Because of this, we spent our time drafting and implementing a new version of the engine while using what we knew about the original engine to maintain the broad strokes of what made it sound unique. The previous vocoder technology was deemed unusable and has taken a significant portion of our development time. We did this rather than use off-the-shelf solutions like HifiGAN to ensure our engine has a unique sound profile that lets it stand out. The articulation engine will be inspired by the first version of the engine, with heavy optimizations and newer module choices.

The current vocoder implementation is still under the development and optimization stage. While it is nearing completion, experimentation and optimization requires compute we haven't had until we recently obtained our new rig. The current vocoder works on lower resolutions, and we're currently trying to scale it to higher ones. As soon as we have decent output, we will post it. Please note that a complete vocoder is not the entire engine; however, the completion of a vocoder will rapidly accelerate the development of the articulation engine.

Languages weren't the hold-up on our development at almost any point. After a while, we decided that it would do the community a disservice to release a language (especially a language without previous commercial support) without a native speaker to represent said language first. The only exceptions to this rule are Oliver, Aurum, and Audine.

We have ten languages fully integrated into our engine: English, French, Portuguese, German, Japanese, Korean, Russian, Spanish, Mandarin Chinese, and Polish. We do not plan to expand this right now and would rather focus our efforts on completing the core engine.

As soon as the engine and editor are in a suitable state, we will be sending out a private beta. At this time, we will review the records for those who bought access to the private beta and extend an invitation to them. We are still under deliberation regarding a public beta.

Current work rebuilding the engine has yielded satisfactory results for the short development window of under a year. It's hard to nail down a timeline right now, but with time, the engine will be something that people will be able to use, with many of the important core features intact.

Our next update will include pictures of the new UI along with a dev blog by Aku, our frontend developer.

The team is self-sufficient, but any relevant volunteer information would be posted in our upcoming Discord server.

Currently: RVBY, Arctos, and Imani, with more planned.

Ongoing product development includes collaborations with SoundCake Studio, Vrilliant, VirVox, and Altvox.