Pev: A Calibrated Fast-Decision Model for Personal Agents
Personalized agents, assistants that keep a long-term memory of one user and act on that user's behalf across shopping, travel, email and calendar, are quickly becoming a product category. Their bottleneck is less open-ended generation than personal judgement: deciding, many times per request, what this particular user would want. Asked to “re-order the dog food”, an agent has to infer that the user switched brands after an allergy, that the new $54 bag now exceeds the user's $50 approval rule, that an address the user asked it to forget must not be reused, and whether the request is worth interrupting the user for. We cast these judgements as seven typed question families and train Pev-27B, a fast decision model that answers each with a calibrated probability from a single forward pass, and release Pev-Bench, built from real public behaviour with user-disjoint splits. On a public test set of 720 new users, Pev-27B reaches 0.915 family-macro accuracy, ahead of gpt-6-astra (0.873), Kev-27B (0.823), Jev (0.788) and its base model (0.762).
- Pev-27B0.915
- gpt-6-astra0.873
- Kev-27B0.823
- Jev0.788
- Qwen3.8-27B (base)0.762

