An AI email assistant needs to do more than write convincing sentences. When a new email arrives, Fyxer first determines whether it needs a reply, a scheduling action, or simply a notification to the user. If a reply is needed, it also analyzes the email’s intent and the expected outcome of the interaction. The important shift for users is that the system is designed to decide whether to draft a reply before deciding what that reply should say. Source
This case is especially relevant to teams building AI products that need to account for email and workplace conversation history, relationships with correspondents, and each person’s writing style. For everyday users, the key is to review an AI-generated draft rather than treat it as a finished reply, checking whether it reflects the earlier conversation and their own intent. The steps below describe Fyxer’s product development approach, not settings readers can simply turn on in their own email service. Source
Before You Start: Break Down the Job
Fyxer did not treat email writing as a single generation task. It separated decisions about whether a reply is needed, intent analysis, retrieval of relevant context, drafting, and other tasks, assigning them to 30–50 specialized models. For teams building workplace AI, the starting point illustrated by this case is to break a broad goal like “write good emails” into smaller tasks. That makes it possible to examine a problem with the draft’s wording separately from a mistake in deciding which emails need replies. Source
Step 1: Decide Which Emails Need Replies
Not every new email follows the same path. Fyxer’s classification stage determines whether a message calls for a reply, a scheduling action, or a simple notification. Only after it determines that a reply is needed do additional models analyze the sender’s intent and the expected outcome of the interaction. For product teams, that is why it matters to assess how an email was classified, not just the quality of the draft. Source
Step 2: Find Relevant Context in Earlier Conversations
The information needed for a reply is not always in the latest email. For each conversation, Fyxer decides what information should be remembered long term, then searches stored interactions for memories relevant to the user and the conversation. The point is not to pull in as much history as possible, but to find what matters to the current exchange. The source does not explain in detail what is stored, how those decisions are made, or how personal information is handled, so those practices cannot be assessed here. Source
Step 3: Build and Validate Against Real Workflows
Before developing the product, Fyxer operated a human executive assistant service and accumulated more than 500,000 hours of annotated workflow data. Those examples covered when to respond, which past conversations were relevant, and how replies varied by user. For model training, Fyxer said it used supervised fine-tuning (SFT) and LoRA to create task-specific variants, and used OpenAI’s fine-tuning platform for tasks where accuracy was especially important. The process illustrates the importance of gathering real workflow examples, not just selecting a model. Source
Before deployment, Fyxer evaluates drafting, classification, and prioritization against its own email-task validation sets. It considers response time and cost alongside accuracy. Teams building a similar product should therefore look beyond whether a reply sounds natural and establish a process for checking task-level results, speed, and cost together. Source
Step 4: Improve Using Users’ Edits
One way to assess whether a draft helped is to compare it with the email the user ultimately sent. Fyxer uses the differences between its original drafts and users’ edited final versions as preference data, creating DPO training data by comparing the two outputs. It says every change to drafting goes through A/B testing and is released only when the improvement is statistically significant. That improvement cycle requires a feedback path for comparing drafts with final versions and an experimental process for validating changes. Source
What to Check When Reviewing the Results
In the Fyxer case presented by OpenAI, 53% of AI-generated drafts were accepted without edits. A co-founder said that, 90 days after signing up, more than 90% of users remained paying customers and used the product daily. The supplied source does not provide the methodology behind those figures or independent verification. They are best read as results reported for this case, not as evidence that another product would achieve the same outcomes. Source
How Far Can You Apply This Example?
If you use AI-generated email drafts, you can check not only whether the tone sounds natural, but also whether the email needs a reply and whether relevant earlier conversations are reflected. For teams building workplace AI, the actionable items illustrated by this case are to examine task decomposition, real-world examples, accuracy, speed, cost, user edits, and A/B testing in sequence. The source does not disclose the detailed configuration or cost of each model, its personal-information practices, or the conditions required to reproduce the results in another product. Source