Banking77
Understand the message. Find the intent.
Fine-tuning a transformer to classify banking queries into 77 intents, evaluated against a baseline.

The problem
A short banking query can express a very specific intent. Telling closely related categories apart requires more than keywords—and a clear understanding of where the model fails.
What I built
- Fine-tuned DistilRoBERTa for the 77 categories in the Banking77 dataset.
- Compared results with TF-IDF and logistic regression, keeping validation and test separate.
- Built a FastAPI demo with CPU inference on Modal for new English queries.
How it works
- New English query
- Tokenization
- Fine-tuned DistilRoBERTa
- One of 77 intents
Engineering decisions
A baseline before a larger model.
TF-IDF and logistic regression provide a useful performance reference. A transformer should justify its added complexity through a measured comparison.
Evaluate without tuning on the test.
Validation was split from the training data. The official test contains 3,080 examples, 40 per category, and was not used to choose parameters.
Separate research from inference.
Falcon-7B-Instruct was explored in the notebook for error analysis. It is not part of the web service: the demo runs only the classifier.
Results and evidence
DistilRoBERTa achieved 92.82% accuracy and 92.82% macro-F1 on the official test. The baseline achieved 85.03% accuracy. Saved metrics and the executed notebook are available in the repository.
View sourceScope and learnings
These are academic Banking77 results, not a guarantee for arbitrary messages. The service requires an access code, may cold-start and depends on available credit. Personal banking information should not be submitted.