An AI email assistant needs to do more than write convincing sentences. When a new email arrives, Fyxer first determines whether it needs a reply, a scheduling action, or simply a notification to the user. If a reply is needed, it also analyzes the email’s intent and the expected outcome of the interaction. The important shift for users is that the system is designed to decide whether to draft a reply before deciding what that reply should say. Source

This case is especially relevant to teams building AI products that need to account for email and workplace conversation history, relationships with correspondents, and each person’s writing style. For everyday users, the key is to review an AI-generated draft rather than treat it as a finished reply, checking whether it reflects the earlier conversation and their own intent. The steps below describe Fyxer’s product development approach, not settings readers can simply turn on in their own email service. Source

Before You Start: Break Down the Job

Fyxer did not treat email writing as a single generation task. It separated decisions about whether a reply is needed, intent analysis, retrieval of relevant context, drafting, and other tasks, assigning them to 30–50 specialized models. For teams building workplace AI, the starting point illustrated by this case is to break a broad goal like “write good emails” into smaller tasks. That makes it possible to examine a problem with the draft’s wording separately from a mistake in deciding which emails need replies. Source

Step 1: Decide Which Emails Need Replies

Not every new email follows the same path. Fyxer’s classification stage determines whether a message calls for a reply, a scheduling action, or a simple notification. Only after it determines that a reply is needed do additional models analyze the sender’s intent and the expected outcome of the interaction. For product teams, that is why it matters to assess how an email was classified, not just the quality of the draft. Source

Step 2: Find Relevant Context in Earlier Conversations

The information needed for a reply is not always in the latest email. For each conversation, Fyxer decides what information should be remembered long term, then searches stored interactions for memories relevant to the user and the conversation. The point is not to pull in as much history as possible, but to find what matters to the current exchange. The source does not explain in detail what is stored, how those decisions are made, or how personal information is handled, so those practices cannot be assessed here. Source

Step 3: Build and Validate Against Real Workflows

Before developing the product, Fyxer operated a human executive assistant service and accumulated more than 500,000 hours of annotated workflow data. Those examples covered when to respond, which past conversations were relevant, and how replies varied by user. For model training, Fyxer said it used supervised fine-tuning (SFT) and LoRA to create task-specific variants, and used OpenAI’s fine-tuning platform for tasks where accuracy was especially important. The process illustrates the importance of gathering real workflow examples, not just selecting a model. Source

Before deployment, Fyxer evaluates drafting, classification, and prioritization against its own email-task validation sets. It considers response time and cost alongside accuracy. Teams building a similar product should therefore look beyond whether a reply sounds natural and establish a process for checking task-level results, speed, and cost together. Source

Step 4: Improve Using Users’ Edits

One way to assess whether a draft helped is to compare it with the email the user ultimately sent. Fyxer uses the differences between its original drafts and users’ edited final versions as preference data, creating DPO training data by comparing the two outputs. It says every change to drafting goes through A/B testing and is released only when the improvement is statistically significant. That improvement cycle requires a feedback path for comparing drafts with final versions and an experimental process for validating changes. Source

What to Check When Reviewing the Results

In the Fyxer case presented by OpenAI, 53% of AI-generated drafts were accepted without edits. A co-founder said that, 90 days after signing up, more than 90% of users remained paying customers and used the product daily. The supplied source does not provide the methodology behind those figures or independent verification. They are best read as results reported for this case, not as evidence that another product would achieve the same outcomes. Source

How Far Can You Apply This Example?

If you use AI-generated email drafts, you can check not only whether the tone sounds natural, but also whether the email needs a reply and whether relevant earlier conversations are reflected. For teams building workplace AI, the actionable items illustrated by this case are to examine task decomposition, real-world examples, accuracy, speed, cost, user edits, and A/B testing in sequence. The source does not disclose the detailed configuration or cost of each model, its personal-information practices, or the conditions required to reproduce the results in another product. Source