What we test, how we test it, and which numbers we deliberately do not publish. A benchmark page that invents its results would be worse than no benchmark page at all.
Tally Assistant publishes what it tests and how, not invented accuracy percentages. Six areas are tested: CSV format parsing (coverage grows with real user uploads, currently 200+ formats), receipt OCR (tested per release against a language sample set), categorization (AI suggests, user confirms by design), currency conversion (daily ECB rates, mathematically verifiable), invoice generation (rule-checked totals in every currency), and client matching (suggested and confirmed). Accuracy percentages for AI features are not published because the models are non-deterministic, and third-party audit results do not exist yet, so none are claimed.
This page publishes test dimensions and methods, not AI accuracy percentages. Those require a fixed model version, fixed dataset, and re-runnable tests we do not have yet.
Every bank and payment-platform layout that parses correctly becomes part of the supported set. Coverage grows continuously as users upload real files, because real files are the only honest test corpus. Each release re-tests the full set.
Published: the supported-format count (200+) and the list of banks with guides. Not published: a percentage accuracy figure.
The scanner is tested per release against a sample set of receipts and payment screenshots in each of the 50+ supported languages, checking that amount, date, and merchant are extracted.
Published: the language coverage and the supported screenshot platforms. Not published: extraction accuracy, because OCR runs through an external model whose output varies per image.
AI suggests a category for every imported transaction, and the user reviews and confirms each suggestion before it is final. That review step is part of the design, not a workaround for a flaw.
Published: the fact that categorization requires a review pass. Not published: an automatic accuracy rate, because the number would be measured before the design's own checkpoint.
Conversion uses daily ECB reference rates from the Frankfurter API. Every currency pair is mathematically verifiable: the transaction keeps its original currency, and reports convert at the rate for that date.
Published: the rate source and that rates refresh daily. You can verify any conversion against the ECB's published rates.
Line items, tax rates, and multi-currency totals follow configured rules. Tests check that a stated amount, a stated tax rate, and a stated currency produce exactly the stated total, in every supported currency.
Published: supported currencies and tax handling. You can verify any invoice total against your own calculation, and the free invoice generator makes that easy.
Transactions are matched to clients from the columns in the file and from payment-history patterns. Matches are suggestions the user confirms, like categorization.
Published: the fact that matching is suggested and confirmable. Not published: a match rate, for the same reason as categorization.
Every published number on this site has a verify path. The methodology page states each one. The shortest version: export a file from your own bank, upload it on the free plan, and compare the result with your statement. That is the test that matters for you.
See how every number is measuredWe publish the supported-format count and the way coverage is tested, but not an accuracy percentage. Parsing runs through an AI model that is non-deterministic and updated regularly, so a fixed percentage would be a false promise. The practical check is yours: upload your own bank file on the free plan and see whether the columns and categories come out right.
Because publishing numbers we cannot defend would violate the methodology we commit to on this site. What we do publish is exactly what we test, how we test it, and how you can verify each claim yourself. That is more useful to a person deciding whether to trust us than a made-up 98%.
When the tests are worth publishing. OCR and categorization run on external models, so meaningful numbers would need a fixed model version, a fixed dataset, and a defined test window, plus a way to re-run them. We are not there yet, and this page will be updated the day we are.
Yes, and the methodology page explains how for every published number. The free plan imports 100 transactions a month, so you can run a real month of your own data through the system and spot-check the results against your bank statement.