← All posts

Can Jev Make Large-Scale Chat History Analysis Practical?

24 September 2026 9 min read Experiments #Jev#TypeSafe#Classification

Jev is new, and I have been experimenting with where this new model class could create genuinely new workflows rather than simply reproduce things we already do with generative LLMs.

The first use case that came to mind was chat-history analysis.

Products accumulate enormous amounts of conversational data: customer-support chats, assistant sessions, sales conversations, internal copilots, community messages and agent traces. The questions we want to ask are straightforward: What are people talking about? Which topics are growing? What percentage of conversations relate to billing, onboarding, errors, cancellations or specific product features? How does the distribution change over time?

The problem is scale.

With today's general-purpose LLMs, analyzing a large conversation archive message by message is possible, but it quickly becomes an inference problem: many calls, substantial token volume, generation latency and cost. Jev is interesting precisely because TypeSafe designed it for a different computation pattern: typed probabilistic decisions, parallel outputs and high-throughput structured classification. TypeSafe currently publishes an input price of $0.042 per million tokens, free output tokens, and 70–500 ms end-to-end response times for System One calls; its workflow benchmarks report gains as high as 193.6× faster and 444.6× cheaper on the evaluated workflows.8

So my question is not simply whether Jev can classify text.

Can Jev turn a large chat history into structured topic data fast and cheaply enough that analysis which is cumbersome with generative LLMs becomes practical?

This first experiment tests that idea on three dimensions at once: speed, economics, and semantic quality.

In this run, the complete 231-message dataset was classified in seconds. The semantic side is measurable because every message has a known ground-truth topic. The economic side is unusually attractive because Jev charges only for input tokens and does not generate an expensive text response for every classification.8

The remaining question is whether that speed and cost come with enough quality to make the resulting topic distribution useful.

Why BANKING77?

BANKING77 was introduced by Casanueva et al. as a challenging single-domain intent-detection dataset containing 13,083 annotated queries across 77 banking intents.1

The number of classes makes the task interesting. Many labels are semantically close: a transfer may be pending, failed, declined, cancelled, charged a fee, or not received; a card payment may be reversed, duplicated, declined, unrecognized, or processed at the wrong exchange rate.

That means the problem is not just keyword matching. The classifier must distinguish between neighboring operational meanings.

BANKING77 is also not a perfect ground truth. Ying and Thomas later reported substantial potential label noise in its training split and showed that removing suspected errors changed downstream classification performance.2 That is another reason to treat a small experiment like this as a diagnostic pilot rather than a definitive leaderboard claim.

A deliberately minimal zero-shot setup

I intentionally avoided elaborate prompt engineering.

The experiment used:

  • 77 intents
  • 3 test examples per intent
  • 231 messages in total
  • zero-shot classification
  • no fine-tuning
  • no few-shot demonstrations
  • no retrieval layer
  • only minimal, human-readable descriptions of the intent labels

For example, the label card_payment_not_recognised was represented as: “The user's intent is about card payment not recognised.”

The instruction was simply to choose exactly one intent and select the most specific class matching the customer's request.

The purpose was not to reproduce the full BANKING77 benchmark protocol. The full test split contains 3,080 examples; our 231-example stratified sample is a pilot. The question was whether Jev could create a useful decision signal with very little task-specific scaffolding.

Result: 79.65% accuracy, 77.83% Macro-F1

Jev classified 184 of 231 messages correctly.

MetricResult
Accuracy79.65%
Macro-F177.83%
Correct predictions184 / 231
Intent classes77

The score is interesting, but the more useful result appeared when we looked at confidence.

Confidence provides a second useful signal

Prediction groupCountMean confidenceMedian confidence
Correct18492.03%99%
Incorrect4774.36%73%

Correct predictions had a mean confidence roughly 17.7 percentage points higher than incorrect predictions.

That does not prove that the confidence score is perfectly calibrated. Calibration has a stricter meaning: a prediction issued with 90% confidence should be correct approximately 90% of the time over a suitable population. Neural classifiers are often miscalibrated, so this must be measured rather than assumed.4

What the pilot does show is that confidence contained useful ranking information: higher-confidence groups were substantially more accurate in this sample.

A second finding: confidence creates an accuracy–coverage control

Instead of accepting every prediction, I applied confidence thresholds.

Minimum confidenceCoverageAccuracyAccepted examples
0.5094.81%81.74%219
0.7082.68%85.86%191
0.8076.19%88.07%176
0.9067.10%91.61%155
0.9560.61%91.43%140

At a 0.90 confidence threshold, Jev retained 155 of the 231 cases—about 67.1% coverage—and accuracy within that accepted subset rose to 91.61%.

This is almost exactly the production question studied by selective classification: how much coverage are we willing to give up in exchange for lower error on the cases the system chooses to handle?3

The slight decrease from 91.61% at 0.90 to 91.43% at 0.95 should also prevent overinterpretation. With only 231 examples, local non-monotonicity is unsurprising. A deployment threshold should be estimated on a larger held-out calibration set, ideally with uncertainty intervals and per-class analysis.4

How this becomes a chat-history analytics pipeline

The most interesting production pattern is not a chatbot or an autonomous agent. It is a semantic map-reduce pipeline over conversational data.

  1. Split the archive into units: messages, turns, sessions, or conversation windows.
  2. Ask Jev to assign each unit to one topic from a predefined taxonomy.
  3. Aggregate those typed decisions into topic distributions, trends, cohorts, and dashboards.
  4. Use confidence to isolate uncertain cases instead of forcing every message into the aggregate.
  5. Send only low-confidence or strategically important cases to a stronger generative LLM or human review.

This is where the speed and cost profile matters. If every message requires a comparatively expensive generative completion, exhaustive analysis becomes difficult as the archive grows. A fast decision model changes the economics because the expensive model becomes an exception path rather than the default path.

The confidence result makes this architecture more practical. It resembles selective classification, where a model can abstain from uncertain examples to improve reliability on the accepted subset.3 It also connects to learning to defer, where difficult cases are routed to another expert.5

A similar economic principle appears in LLM cascading work such as FrugalGPT and RouteLLM: use cheaper computation for the cases it can handle and reserve more expensive models for the cases that need them.67

For chat-history analytics, however, the primary goal is simpler: convert a very large unstructured conversation archive into structured, measurable topic data fast enough and cheaply enough to analyze all of it—not merely a sample.

Accuracy alone hides two very different systems

Consider two deployments:

System A: classify 100% of requests at 79.65% accuracy.

System B: automatically classify roughly 67% at 91.61% accuracy and escalate the rest.

They share the same underlying classifier, but their operational risk is very different.

For a low-impact workflow, full coverage may be acceptable. For banking, legal, healthcare, security, or other high-impact processes, selective automation can be more attractive because uncertainty becomes an explicit system state rather than an invisible model failure.

This is why coverage, calibration, and deferral deserve to sit next to accuracy on an evaluation dashboard.345

What I would test next

1. Run the full 3,080-example test split.

The pilot is too small for strong per-intent conclusions.

2. Measure calibration explicitly.

Reliability diagrams, Expected Calibration Error, Brier-style metrics, and class-conditional calibration would tell us whether Jev's confidence is numerically trustworthy, not merely correlated with correctness.4

3. Improve only ambiguous label descriptions.

Rather than prompt-engineer all 77 classes, identify confusion clusters and enrich only the boundaries that need more semantic information.

4. Compare routing architectures.

Test Jev-only, strong-LLM-only, and Jev → LLM cascade designs on accuracy, latency, tokens, and cost per successful task.67

5. Add human deferral for high-risk cases.

The relevant metric then becomes system-level error and workload allocation, not model accuracy in isolation.5

Takeaway

I started this experiment with a use-case question rather than a benchmark question:

Can Jev make large-scale chat-history analysis fast and inexpensive enough to become a practical analytics primitive?

The first result is encouraging.

On a proxy made of 231 conversation-like messages spread across 77 closely related topics, Jev processed the complete dataset in seconds while reaching 79.65% accuracy and 77.83% Macro-F1 with minimal category descriptions and no task-specific training examples.

That combination matters more than any single metric.

A slower generative model may achieve higher accuracy. A cheap classifier may be fast but semantically weak. What makes Jev interesting here is the combination of speed, extremely low inference economics, high-cardinality structured decisions, and usable semantic accuracy.

Confidence adds a second useful property. Correct predictions averaged 92.03% confidence versus 74.36% for incorrect predictions, and a 0.90 confidence threshold produced a 91.61%-accurate subset at 67.10% coverage. This suggests that uncertain classifications can be separated from the main analytical pipeline rather than silently contaminating the aggregate.

So my current conclusion is:

Jev looks like a serious candidate for turning very large volumes of chat data into structured topic distributions in seconds and at a cost profile that changes what is practical to analyze.

The next step is no longer to ask whether the idea works at all. It is to stress-test it on much larger real chat histories and measure how throughput, total cost, topic-distribution error and confidence calibration behave as the dataset grows.

References

  1. Casanueva, I., Temčinas, T., Gerz, D., Henderson, M., & Vulić, I. (2020). Efficient Intent Detection with Dual Sentence Encoders. NLP4ConvAI / ACL. Introduces BANKING77: 13,083 annotated examples across 77 banking intents. Paper
  2. Ying, C., & Thomas, S. (2022). Label Errors in BANKING77. ACL Workshop on Insights from Negative Results in NLP. Reports potential label noise in the BANKING77 training set and shows its effect on classification results. Paper
  3. Geifman, Y., & El-Yaniv, R. (2017). Selective Classification for Deep Neural Networks. NeurIPS 2017. Formalizes the accuracy/risk-versus-coverage trade-off when a classifier is allowed to reject uncertain cases. Paper
  4. Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On Calibration of Modern Neural Networks. ICML 2017. Shows why model confidence should not automatically be interpreted as empirical correctness probability and studies post-hoc calibration. Paper
  5. Mozannar, H., & Sontag, D. (2020). Consistent Estimators for Learning to Defer to an Expert. ICML 2020. Studies systems that either make a prediction or defer a case to a downstream expert. Paper
  6. Chen, L., Zaharia, M., & Zou, J. (2023). FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. Describes LLM cascades that route queries across models to trade off quality and inference cost. Paper
  7. Ong, I. et al. (2024). RouteLLM: Learning to Route LLMs with Preference Data. Studies learned routers between stronger and weaker LLMs to reduce cost while preserving response quality. Paper
  8. TypeSafe AI (2026). Introducing System One Models & Jev. Official product/research announcement describing Jev, typed probabilistic decisions, confidence, and the RLCD framing. This is a company source, not a peer-reviewed paper. Source
  9. TypeSafe AI (2026). System One API documentation. Official API reference for Choice, Score, Noul and Jev model endpoints. This is product documentation, not an academic paper. Source