SMILE (haifengl/smile): A JVM Machine Learning Framework With a Java 25 Requirement
Statistical Machine Intelligence & Learning Engine
At a glance
- What is it?
- SMILE is a statistical machine learning library for the JVM with Java, Scala and Kotlin APIs. Its v5+ line requires Java 25, which is the first thing to check before adopting it.
- Who is it for?
- Adopt SMILE if your stack is already on the JVM and you want classification, regression, clustering, NLP and numerical routines behind one dependency instead of stitching several libraries together. Skip it if you are pinned to Java 21 or Java 8, since the README states that v5+ requires Java 25 and older versions require Java 8, and skip it if you want a Python-first workflow.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SMILE Solves for JVM Teams
Java has never lacked numerical code, but it has lacked a single coherent place to put it. A typical JVM data project ends up with one dependency for linear algebra, another for a decision tree, a third for CSV parsing, and a fourth for plotting. SMILE's answer is to ship all of that in one repository with a shared data representation and a shared module layout.
The scope listed in the README is broad: classification (SVM, random forest, AdaBoost, gradient boosting, logistic regression, KNN, naive Bayes, LDA/QDA/RDA), regression (SVR, Gaussian process, LASSO, ElasticNet, ridge), clustering (k-means, DBSCAN, BIRCH, spectral, SOM), manifold learning (t-SNE, UMAP, IsoMap, LLE, PCA), feature engineering, NLP, association rule mining, sequence learning, nearest neighbor search, numerical optimization, and visualization.
The audience is specific. This is for engineers and analysts who already deploy JVM services and want models in the same process and the same build, rather than exporting a pickle file from a Python training job and calling it over HTTP. The README also states that SMILE provides idiomatic APIs for Scala and Kotlin, so polyglot JVM teams are explicitly in scope.
Modules, the DataFrame Layer, and the LLM Addition
The repository is organized as a set of top-level modules rather than one monolith. The README's module map names base/ for data structures, math, linear algebra and I/O, and core/ for the learning algorithms. Around those sit deep/ (LibTorch and GPU backend), nlp/, plot/, json/, spark/, scala/, kotlin/, clojure/, serve/ and studio/.
base/ is where the data flow starts. The README points to DATA_FRAME.md for a DataFrame API covering creation, selection and transformation, and to DATA_IO.md for readers and writers across CSV, JSON, Parquet, Arrow, JDBC and Avro. That is the practical center of the library: data lands in a DataFrame, transformations and a formula API build a model matrix, and core/ algorithms consume it.
The most recent direction is LLM support. The feature table lists LLaMA-3 inference, a tiktoken BPE tokenizer, an OpenAI-compatible REST server and SSE chat streaming, with the serve/ module in the repository tree. This is a notable expansion for a library whose history is statistical learning, and it means the dependency surface is no longer only numerical. If you only want a random forest, you are still pulling in a project that now also ships inference and serving code.
SMILE Studio is described in the README as an agentic IDE for data science using Python, Java or Scala, with instructions in studio/README.md. Treat that as a separate product with its own setup path, not as part of the library install.
Install Routes and the First Real Project
The README documents four installation routes: Maven, SBT for Scala, Gradle for Kotlin, and native libraries for BLAS/LAPACK. The Maven badge links to the com.github.haifengl/smile-core artifact on Maven Central, and the module map shows that base/ holds the foundation while core/ holds the learning algorithms, so a project using both resolves both artifacts.
The README's installation section does not print a dependency snippet or any command line, so there is no code to copy here. The only concrete install details it gives are the build tool names above and the artifact named in the badge; the version string has to come from Maven Central rather than from an article.
For Scala the README documents SBT, and for Kotlin it documents Gradle. Those resolve the same artifacts through a different build tool; they are not separate libraries.
The native libraries section matters more than it looks. SMILE's linear algebra can run against native BLAS and LAPACK implementations, and the README devotes a subsection to it. On a laptop the pure-Java path is usually fine; in a container or on a cluster node, the native path is a deployment step rather than a code change.
Before any of this, confirm your JDK. The README states that SMILE v5+ requires Java 25, v4.x requires Java 21, and all previous versions require Java 8. That single sentence determines which version line you can use.
Once the dependency resolves, the workflow the README describes is DataFrame in, model out. Data is read through the base/ I/O layer, transformed, and passed to an algorithm in core/. The README points to core/TRAINING.md for cross-validation, bootstrap and hyper-parameter search, and to core/VALIDATION.md and core/VALIDATION_METRICS.md for hold-out, k-fold and leave-one-out evaluation plus accuracy, AUC, F1, RMSE, MAE and confusion matrix reporting. Those three documents are where a first real project should go next, because they cover the part that turns a model into a defensible result.
The Java 25 Floor and Other Adoption Costs
The version-to-JDK mapping is the hardest constraint in the project. If your organization standardizes on Java 21, you are on the v4.x line, not the current one, and you do not get whatever landed in v5 and later. The README does not describe a backport policy for the newer line, so planning around an older JDK means planning around an older feature set.
The second cost is dependency weight. Choosing SMILE for a classifier also brings in the DataFrame and I/O layer, the plotting module, the NLP module and the LLM serving code, depending on which artifacts you resolve. Modular artifacts mitigate this, but the module map is large enough that reviewing what you actually pull in is worthwhile before a first production build.
The third is documentation shape. The README is a map, not a manual: it links to roughly thirty topic documents under base/ and core/ alone. That is good for depth and awkward for onboarding, because there is no single tutorial that walks from an empty project to a validated model. Expect to read DATA_FRAME.md, DATA_IO.md and one of the VALIDATION documents before writing much code.
A fourth point is the licence. The repository metadata reports the licence as NOASSERTION, and the repository contains both LICENSE and COPYING files. NOASSERTION means the automated classifier could not map the licence to a known SPDX identifier, not that the project is unlicensed. Anyone embedding SMILE in a distributed product should read those two files directly rather than relying on a badge or a summary.
Where SMILE Is the Wrong Choice
SMILE is the wrong tool when the model is the product and the language is not. If your team writes Python and deploys models as services, the JVM requirement is a tax with no matching benefit, and the ecosystem of Python libraries around gradient boosting and deep learning is deeper.
It is also the wrong tool when you need only one algorithm. Pulling in a framework to fit a single logistic regression is a poor trade against a small dedicated dependency, and the module structure means you will still be reading framework documentation to get there.
There is a sharper limitation in the LLM area. The README lists LLaMA-3 inference and an OpenAI-compatible REST server. That is a specific model family and a specific serving shape. If your requirement is a different model architecture, or a serving stack with batching and scheduling guarantees, the README does not claim to cover it. The LLM support reads as a useful addition for JVM teams that want inference in-process, not as a replacement for a dedicated inference server.
Finally, the visualization story is desktop-oriented: Swing plots plus declarative Vega-Lite charts. Swing is a poor fit for headless server rendering, so if chart output is part of your pipeline, the Vega-Lite path is the one to evaluate, and the README does not document a server-side rendering route.
SMILE Against scikit-learn and Weka
The closest comparison for a Python user is scikit-learn. The difference is not the algorithm list, which overlaps heavily, but the runtime and the data layer. scikit-learn sits on NumPy and pandas; SMILE sits on its own DataFrame and tensor types and runs on the JVM. If your data already lives in a JVM service, SMILE removes a serialization boundary. If it lives in a notebook, scikit-learn removes a language boundary. Neither is strictly better; they optimize for different starting points.
Against Weka, the difference is intent. Weka is a workbench with a GUI and an explorer-driven workflow, built for experimentation and teaching. SMILE is a library you call from code, with the module map organized around APIs rather than an application. If you want to click through datasets and compare classifiers interactively, Weka fits that. If you want a classifier inside a Spring service, SMILE fits that.
There is a third comparison inside the JVM ecosystem worth naming: for a single algorithm, a small focused library is often the better dependency. SMILE's advantage appears when you need several families at once, for example clustering to segment, a classifier to score, and a validation framework to report both, sharing one DataFrame implementation.
Release Cadence and Upgrade Planning
The release history shown here is recent and steady: v6.3.0 on 2026-08-18, v6.2.5 on 2026-08-02, and v6.2.4 on 2026-07-13. The repository is not archived, and the last push was on 2026-09-09. That pattern suggests patch releases arrive within weeks and minor releases within a couple of months, so pinning an exact version in your build file and upgrading deliberately is more realistic than tracking the latest.
The upgrade cost concentrates in two places. The first is the JDK floor: a jump from the v4.x line to v5+ is a Java 21 to Java 25 move, which is an infrastructure change, not a dependency bump. The second is the breadth of the library itself. Because base/, core/, deep/, nlp/ and serve/ evolve together, a minor release can touch code you do not use but still compile against. Reading the release notes before upgrading is the cheap mitigation, and the README's module-level documents are the place to check whether an API you depend on moved.
On licensing, the only defensible statement is procedural: the repository carries LICENSE and COPYING files and the metadata reports NOASSERTION. Read both files and confirm the terms with whoever owns compliance in your organization. Nothing in the README substitutes for that.
Editorial conclusion
Adopt SMILE if your stack is already on the JVM and you want classification, regression, clustering, NLP and numerical routines behind one dependency instead of stitching several libraries together. Skip it if you are pinned to Java 21 or Java 8, since the README states that v5+ requires Java 25 and older versions require Java 8, and skip it if you want a Python-first workflow. Before committing, verify three things: which SMILE major version matches your JDK, whether the native BLAS/LAPACK libraries are available on your target machines, and what the LICENSE and COPYING files actually say, because the repository metadata reports the licence as NOASSERTION rather than a recognized SPDX identifier.
Frequently asked questions
What Java version does SMILE require?
The README states that SMILE v5+ requires Java 25, v4.x requires Java 21, and all previous versions require Java 8. Which SMILE version you can use is therefore determined by the JDK your project runs on.
How do I install SMILE with Maven?
The README documents Maven, SBT for Scala, and Gradle for Kotlin as installation routes. The Maven badge points to the com.github.haifengl/smile-core artifact on Maven Central, and the module map shows base/ and core/ as separate modules.
Does SMILE support deep learning and LLM inference?
The feature table lists a LibTorch/GPU backend, EfficientNet-V2 image classification and a custom layer API under deep learning, and LLaMA-3 inference, a tiktoken BPE tokenizer, an OpenAI-compatible REST server and SSE chat streaming under LLM.
What licence does SMILE use?
The repository metadata reports the licence as NOASSERTION, and the repository contains both LICENSE and COPYING files. That means the automated classifier could not match it to a known SPDX identifier, so the files themselves are the source to read.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/haifengl-smile)