{"slug": "ml-net", "title": "ML.NET", "summary": "A deep-dive walkthrough documents ML.NET, Microsoft's open-source, cross-platform machine learning framework that lets .NET developers build, train, and run models entirely in C# or F# without leaving the .NET ecosystem. The guide covers the full workflow — loading data into an IDataView, building estimator pipelines, fitting and evaluating models, saving and loading them, and serving predictions in production — plus AutoML and worked examples for classification, regression, clustering, recommendation, and anomaly detection. It frames ML.NET as a pragmatic choice for classic ML on tabular and text data, while recommending a train-in-Python, export-to-ONNX pattern for deep learning.", "body_md": "*A deep-dive walkthrough of ML.NET — Microsoft's open-source, cross-platform machine learning framework for .NET — covering `MLContext`, data loading and `IDataView`, preprocessing and feature engineering, how training pipelines (estimators and transformers) actually work, model training and evaluation, saving/loading models, making predictions safely in production, AutoML, and complete worked examples for classification, regression, clustering, recommendation systems, and anomaly detection.*\n\nML.NET lets .NET developers build, train, and run machine learning models **entirely in C# or F#**, without switching to Python or leaving the .NET ecosystem. A model trained with ML.NET is just a file you load into an ordinary .NET application — a web API, a worker service, a desktop app — and call like any other dependency.\n\n``` js\nvar mlContext = new MLContext(seed: 0);\n\nIDataView data = mlContext.Data.LoadFromTextFile<HouseData>(\"houses.csv\", hasHeader: true, separatorChar: ',');\n\nvar pipeline = mlContext.Transforms.Concatenate(\"Features\", \"Size\", \"Bedrooms\")\n    .Append(mlContext.Regression.Trainers.Sdca(labelColumnName: \"Price\"));\n\nITransformer model = pipeline.Fit(data);             // TRAINING\nvar engine = mlContext.Model.CreatePredictionEngine<HouseData, HousePrediction>(model);\nvar result = engine.Predict(new HouseData { Size = 1800, Bedrooms = 3 });   // INFERENCE\n```\n\nThose six lines contain the whole ML.NET workflow: **load data → build a pipeline → fit it → predict.** Everything else in this guide is detail on one of those steps. If the vocabulary here (features, labels, training vs. inference, overfitting, precision/recall) is unfamiliar, this series' *AI/ML Fundamentals* guide covers the concepts; this guide focuses on how to apply them in .NET.\n\nML.NET is a machine learning library for .NET. It provides data loading, transformation, a catalog of training algorithms, evaluation tools, and model persistence, all behind a consistent API. It runs on Windows, Linux, and macOS, and is designed for **integrating ML into .NET applications** rather than for research.\n\n```\nStrong fit:\n  - Classic ML on tabular and text data (classification, regression, clustering,\n    recommendation, anomaly detection, time series)\n  - Teams that are .NET-first and want ML inside their existing services\n  - Training and serving in the same language, with no separate Python service\n  - Using models trained elsewhere via ONNX import, then serving them from .NET\n\nWeaker fit:\n  - Cutting-edge deep learning research and large-scale neural network training\n    (the Python ecosystem — PyTorch, TensorFlow — is far richer here)\n  - Needing the very latest model architectures as soon as they're published\n```\n\nAn honest framing: ML.NET is the pragmatic choice for classic ML inside .NET applications. For deep learning, a common pattern is to train in Python, export to **ONNX**, and consume the model from .NET (ML.NET and ONNX Runtime both support this).\n\n```\nTask                      Example question                          Catalog\n------------------------  ----------------------------------------  ---------------------------\nBinary classification     Is this review positive?                  mlContext.BinaryClassification\nMulticlass classification Which category is this ticket?            mlContext.MulticlassClassification\nRegression                What will this house sell for?            mlContext.Regression\nClustering                Which customers behave alike?             mlContext.Clustering\nRecommendation            Which movies will this user like?         mlContext.Recommendation()\nAnomaly detection         Is this data point unusual?               mlContext.AnomalyDetection / Transforms.Detect*\nRanking / Forecasting     Order results; predict future values      mlContext.Ranking / Forecasting\ndotnet add package Microsoft.ML                  # core: data, transforms, common trainers\ndotnet add package Microsoft.ML.AutoML           # AutoML (Section 17)\ndotnet add package Microsoft.ML.Recommender      # matrix factorization (Section 15)\ndotnet add package Microsoft.ML.TimeSeries       # time-series anomaly detection (Section 16)\ndotnet add package Microsoft.ML.FastTree         # gradient-boosted tree trainers\ndotnet add package Microsoft.Extensions.ML       # PredictionEnginePool for ASP.NET Core (Section 11)\n1. Create an MLContext\n2. Load data into an IDataView\n3. Split into train / test sets\n4. Build a pipeline: preprocessing transforms + a trainer\n5. Fit the pipeline on the training set  -> a trained model (ITransformer)\n6. Evaluate on the test set\n7. Save the model to a file\n8. Load it in your application and make predictions\njs\nvar mlContext = new MLContext(seed: 0);\n```\n\n`MLContext` is the factory and the catalog root. Every operation hangs off it:\n\n``` php\nmlContext.Data           -> loading, saving, splitting, filtering data\nmlContext.Transforms     -> preprocessing and feature engineering\nmlContext.Regression     -> regression trainers and evaluation\nmlContext.BinaryClassification / MulticlassClassification\nmlContext.Clustering / AnomalyDetection / Ranking / Forecasting\nmlContext.Recommendation()    -> recommendation trainers\nmlContext.Model          -> save / load / create prediction engines\nmlContext.Auto()         -> AutoML (needs Microsoft.ML.AutoML)\n```\n\n`seed` parameter matters for reproducibility\n\n```\nPassing a seed makes operations that involve randomness (data shuffling in\nsplits, some trainers' initialization) produce repeatable results — which\nis essential when comparing two pipeline variants, or when a teammate needs\nto reproduce your numbers. Without a seed, results can differ slightly run to run.\n```\n\n`MLContext` also exposes logging (`mlContext.Log += ...`) for observing what trainers are doing, and it should generally be **created once and reused** rather than constructed repeatedly.\n\n```\npublic class HouseData\n{\n    [LoadColumn(0)] public float Size { get; set; }\n    [LoadColumn(1)] public float Bedrooms { get; set; }\n    [LoadColumn(2)] public string Neighborhood { get; set; }\n    [LoadColumn(3)] public float Price { get; set; }       // the label we want to predict\n}\n\npublic class HousePrediction\n{\n    [ColumnName(\"Score\")] public float Price { get; set; } // regression trainers write predictions to \"Score\"\n}\nIDataView data = mlContext.Data.LoadFromTextFile<HouseData>(\n    path: \"houses.csv\",\n    hasHeader: true,\n    separatorChar: ',');\n```\n\nNotes worth knowing precisely:\n\n```\n- [LoadColumn(n)] maps a class property to the n-th column in the file\n  (zero-based). [LoadColumn(1, 5)] maps a RANGE of columns into a float[] .\n- Use float, not double, for numeric columns — it's ML.NET's native numeric type.\n- [ColumnName(\"Label\")] renames a column; ML.NET trainers look for a column\n  named \"Label\" by default (and \"Features\" for the feature vector).\n// From an in-memory collection (great for tests and small data)\nIDataView fromList = mlContext.Data.LoadFromEnumerable(houses);\n\n// From a database (requires the Microsoft.ML package's database loader support)\nvar loader = mlContext.Data.CreateDatabaseLoader<HouseData>();\nvar dbSource = new DatabaseSource(SqlClientFactory.Instance, connectionString,\n                                  \"SELECT Size, Bedrooms, Neighborhood, Price FROM Houses\");\nIDataView fromDb = loader.Load(dbSource);\n```\n\nLoading from SQL Server is a natural fit for .NET teams — your training data is often already in a database (see this series' SQL guides). Whichever source you use, the result is the same type: an `IDataView`.\n\n``` js\nvar split = mlContext.Data.TrainTestSplit(data, testFraction: 0.2, seed: 1);\nIDataView trainData = split.TrainSet;\nIDataView testData  = split.TestSet;\nSplit BEFORE any fitting, so the test set never influences training.\nFor time-ordered data, do NOT use a random split — split by time instead\n(train on the past, test on the future), or you'll leak future information.\n```\n\n`IDataView` is the type that flows through every part of ML.NET — what loaders produce, what transforms consume and emit, and what trainers learn from. Think of it as a **read-only, lazily-evaluated table with a schema.**\n\n```\nKey properties:\n  - LAZY: nothing is read or computed until something iterates over it.\n    Building a pipeline doesn't touch the data; only Fit/Transform-and-iterate does.\n  - FORWARD-ONLY cursor access: designed to stream data, so datasets larger\n    than memory can be processed.\n  - SCHEMA-ful: every column has a name and a type (float, string, vector, key...).\n  - IMMUTABLE: transforms produce a NEW IDataView layered on the previous one;\n    the original is never modified.\njs\n// Peek at the first rows and the schema — the standard debugging move\nvar preview = data.Preview(maxRows: 5);\nforeach (var row in preview.RowView)\n    Console.WriteLine(string.Join(\", \", row.Values.Select(kv => $\"{kv.Key}={kv.Value}\")));\n\n// Pull a single column out\nIEnumerable<float> prices = data.GetColumn<float>(\"Price\");\n\n// Convert back to a strongly typed C# collection\nvar houses = mlContext.Data.CreateEnumerable<HouseData>(data, reuseRowObject: false).ToList();\nBecause an IDataView is lazy, a transform that is expensive (text featurization,\nsay) is RE-COMPUTED every time something iterates the data — and iterative\ntrainers iterate many times. Cache after expensive steps:\n\n  pipeline.AppendCacheCheckpoint(mlContext).Append(trainer)\n\n(See Section 7. Skip caching for very large data that won't fit in memory.)\n```\n\nAlso: `reuseRowObject: false` in `CreateEnumerable` is essential if you store the results — with `true`, the same object is overwritten on each iteration, so a collected list would contain many references to one final row's values.\n\nPreprocessing transforms live in `mlContext.Transforms`. Each one is an **estimator** that you chain into a pipeline (Section 7).\n\n```\n// Missing values: replace with the column's mean (default) before training\nmlContext.Transforms.ReplaceMissingValues(\"Size\", replacementMode: MissingValueReplacingEstimator.ReplacementMode.Mean)\n\n// Categorical text -> numbers: one-hot encode\nmlContext.Transforms.Categorical.OneHotEncoding(\"NeighborhoodEncoded\", \"Neighborhood\")\n\n// Scale numeric features to a comparable range\nmlContext.Transforms.NormalizeMinMax(\"Features\")        // squashes to [0, 1]\nmlContext.Transforms.NormalizeMeanVariance(\"Features\")  // zero mean, unit variance\n\n// Map a string label to the \"key\" type multiclass trainers require\nmlContext.Transforms.Conversion.MapValueToKey(\"Label\", \"Category\")\n// Drop obviously bad rows (e.g. impossible prices) BEFORE training\nIDataView cleaned = mlContext.Data.FilterRowsByColumn(\n    data, columnName: \"Price\", lowerBound: 10_000, upperBound: 5_000_000);\n- Models consume NUMBERS. Strings (neighborhood names) must be encoded.\n- Missing values can crash or silently distort a trainer; decide a policy explicitly.\n- Features on wildly different scales (Size ~ 2000, Bedrooms ~ 3) can make\n  gradient-based trainers converge slowly or favor the large-valued feature.\n  Normalization puts them on equal footing.\n- Tree-based trainers (FastTree, LightGBM) are far less sensitive to scale, so\n  normalization matters most for linear models and neural-network-style learners.\n```\n\n`Features` vector trainers consume\nEvery ML.NET trainer expects **one vector column** (named `Features` by default) containing all the model's inputs. `Concatenate` builds it:\n\n```\nmlContext.Transforms.Concatenate(\"Features\", \"Size\", \"Bedrooms\", \"NeighborhoodEncoded\")\nmlContext.Transforms.Text.FeaturizeText(\"Features\", \"ReviewText\")\n```\n\n`FeaturizeText` tokenizes, normalizes, and converts text into a numeric vector of word and character n-gram features — it's the standard starting point for any text classification problem, such as sentiment analysis (Section 12).\n\n``` php\nTransforms.NormalizeBinning(...)           -> bucket a numeric column into ranges\nTransforms.Categorical.OneHotHashEncoding  -> hash-encode HIGH-cardinality categories\n                                              (thousands of distinct values) compactly\nTransforms.Text.ProduceWordBags(...)       -> bag-of-words / n-grams\nTransforms.Conversion.ConvertType(...)     -> change a column's data type\nTransforms.SelectColumns / DropColumns     -> keep only what you need\nTransforms.FeatureSelection.*              -> select the most informative features\nTransforms.CustomMapping(...)              -> your own C# logic (with a caveat below)\n```\n\n`CustomMapping`\n\n```\nCustomMapping lets you run arbitrary C# per row — flexible, but a model that\ncontains one CANNOT simply be saved and loaded in another process unless the\nmapping is registered through a contract assembly. Prefer a built-in transform,\nor do the derived-feature computation in your data-loading code BEFORE ML.NET\nsees the data, so the saved model stays self-contained.\nGood features typically move model quality more than choosing a fancier algorithm:\n  - Derive \"price per square foot\" or \"age of house\" from raw columns\n  - Extract \"day of week\" or \"hour\" from a timestamp\n  - Combine related fields into ratios or flags\nDo this thoughtfully, and avoid features that wouldn't exist at prediction time\n(data leakage — see this series' AI/ML Fundamentals guide).\nIEstimator<T>   — a RECIPE for a step. It hasn't seen data yet.\n                  It may need to LEARN something from data (the mean to\n                  normalize by, the vocabulary to encode, a model's weights).\n\nITransformer    — the RESULT of fitting an estimator. It has learned its\n                  parameters and can now TRANSFORM data.\n\nestimator.Fit(trainingData)  ->  ITransformer\ntransformer.Transform(data)  ->  IDataView with new/changed columns\njs\nvar pipeline = mlContext.Transforms.ReplaceMissingValues(\"Size\")\n    .Append(mlContext.Transforms.Categorical.OneHotEncoding(\"NeighborhoodEncoded\", \"Neighborhood\"))\n    .Append(mlContext.Transforms.Concatenate(\"Features\", \"Size\", \"Bedrooms\", \"NeighborhoodEncoded\"))\n    .Append(mlContext.Transforms.NormalizeMinMax(\"Features\"))\n    .AppendCacheCheckpoint(mlContext)                                    // cache the preprocessed data\n    .Append(mlContext.Regression.Trainers.Sdca(labelColumnName: \"Price\", featureColumnName: \"Features\"));\n\nITransformer model = pipeline.Fit(trainData);   // fits EVERY step, in order, on trainData\n1. ONE object captures preprocessing AND the model. When you save the trained\n   model (Section 10), the normalization parameters, encoders, and trainer all\n   travel together — so predictions in production apply EXACTLY the same\n   transformations as training. This eliminates a classic bug: preprocessing\n   in production subtly differing from preprocessing at training time.\n\n2. Fitting happens on the TRAINING data only. The normalization mean/range is\n   learned from the training set and merely APPLIED to the test set —\n   preventing test-set information from leaking into training.\n\n3. Pipelines are lazy and composable: define once, Fit, Transform, reuse.\n```\n\n`AppendCacheCheckpoint` stores the data computed up to that point in memory so iterative trainers don't recompute upstream transforms on every pass. Place it **after** the expensive preprocessing and **before** the trainer. For datasets too large for memory, omit it.\n\n```\nITransformer model = pipeline.Fit(trainData);\n```\n\n`Fit` runs the whole pipeline over the training data: transforms learn their parameters, and the trainer runs its optimization to learn the model's weights. The returned `ITransformer` is your **trained model** (preprocessing + predictor together).\n\nML.NET's catalog offers many trainers per task. A practical way to choose:\n\n```\nStart simple and fast:\n  - Linear trainers (Sdca*, Lbfgs*, Sgd*): fast, a good baseline, work well on\n    sparse/text data; benefit from normalization.\nThen try stronger:\n  - Tree-ensemble trainers (FastTree*, FastForest*, LightGbm*): often the best\n    accuracy on tabular data; less sensitive to scaling; handle non-linear patterns.\n// Swapping a trainer is a ONE-LINE change — the rest of the pipeline is untouched\n.Append(mlContext.Regression.Trainers.Sdca(labelColumnName: \"Price\"))        // linear baseline\n.Append(mlContext.Regression.Trainers.FastTree(labelColumnName: \"Price\"))    // gradient-boosted trees\n```\n\nThat one-line interchangeability is a major practical strength: you can compare several algorithms against the same preprocessing and the same test set with minimal code (or let AutoML do it, Section 17).\n\n```\nTrainers take labelColumnName and featureColumnName. If your label column is\nliterally \"Label\" and features are in \"Features\", you can omit them. Otherwise\npass them explicitly — a mismatched column name is the most common\n\"Schema mismatch\" error beginners hit.\nIDataView predictions = model.Transform(testData);          // run the trained model on the TEST data\nRegressionMetrics metrics = mlContext.Regression.Evaluate(predictions, labelColumnName: \"Price\");\n\nConsole.WriteLine($\"R²:   {metrics.RSquared:F3}\");\nConsole.WriteLine($\"RMSE: {metrics.RootMeanSquaredError:F0}\");\nConsole.WriteLine($\"MAE:  {metrics.MeanAbsoluteError:F0}\");\n```\n\nEach task has its own evaluator and metrics:\n\n```\nTask                    Evaluate call                                   Key metrics\n----------------------  ----------------------------------------------  -----------------------------------------\nRegression              mlContext.Regression.Evaluate                   RSquared, RootMeanSquaredError, MeanAbsoluteError\nBinary classification   mlContext.BinaryClassification.Evaluate         Accuracy, AreaUnderRocCurve, F1Score,\n                                                                        PositivePrecision, PositiveRecall\nMulticlass              mlContext.MulticlassClassification.Evaluate     MicroAccuracy, MacroAccuracy, LogLoss\nClustering              mlContext.Clustering.Evaluate                   AverageDistance, DaviesBouldinIndex\nAnomaly detection       mlContext.AnomalyDetection.Evaluate             AreaUnderRocCurve, DetectionRateAtFalsePositiveCount\njs\nvar cvResults = mlContext.Regression.CrossValidate(\n    data, pipeline, numberOfFolds: 5, labelColumnName: \"Price\");\n\ndouble avgR2 = cvResults.Average(r => r.Metrics.RSquared);\n```\n\nA single split can be lucky or unlucky; cross-validation trains and evaluates across several folds and lets you look at both the average *and* the spread.\n\n```\n- NEVER evaluate on the training data — it measures memorization.\n- Compare against a baseline; a \"good\" metric means little without one.\n- Look at the train-vs-test gap: great training metrics with poor test metrics\n  means overfitting.\n- On imbalanced classification, don't trust accuracy alone — check precision,\n  recall, F1, and AUC.\n```\n\n`.zip` file you can ship\n\n```\n// Save: the model AND the schema of the data it was trained on\nmlContext.Model.Save(model, trainData.Schema, \"model.zip\");\n\n// Load (in the same app, or — typically — in a different one)\nvar loadMlContext = new MLContext();\nITransformer loadedModel = loadMlContext.Model.Load(\"model.zip\", out DataViewSchema inputSchema);\n```\n\nThe saved file contains the **entire pipeline** — preprocessing steps with their learned parameters, plus the trained predictor — so the consuming application doesn't need the training data or the training code, only the file and the matching C# input/output classes.\n\n```\nPractical guidance:\n  - Treat model.zip as a build/deployment artifact, versioned like any other.\n    Name or folder it by version (model-v3.zip) so a rollback is trivial.\n  - Keep the input/output classes (HouseData, HousePrediction) in a shared\n    project referenced by both the training and the consuming app, so the\n    schema can't drift apart.\n  - Models can also be saved/loaded via Stream — useful for storing them in\n    blob storage or a database rather than on local disk.\n```\n\n`PredictionEngine`\n\n``` js\nvar engine = mlContext.Model.CreatePredictionEngine<HouseData, HousePrediction>(model);\n\nHousePrediction result = engine.Predict(new HouseData { Size = 1800, Bedrooms = 3, Neighborhood = \"Westside\" });\nConsole.WriteLine($\"Predicted price: {result.Price:C0}\");\nA PredictionEngine holds internal state and must not be shared across threads.\nCreating a new one per request is wasteful (it's relatively expensive to build).\nIn a web application, using a single shared PredictionEngine instance — a very\ncommon mistake — produces race conditions and wrong results under concurrent load.\n```\n\n`PredictionEnginePool`\n\n```\n// Program.cs  (requires the Microsoft.Extensions.ML package)\nbuilder.Services.AddPredictionEnginePool<HouseData, HousePrediction>()\n    .FromFile(modelName: \"HouseModel\", filePath: \"model.zip\", watchForChanges: true);\n// In a controller / minimal API endpoint — inject the pool\napp.MapPost(\"/predict\", (HouseData input, PredictionEnginePool<HouseData, HousePrediction> pool) =>\n{\n    var prediction = pool.Predict(modelName: \"HouseModel\", example: input);\n    return Results.Ok(prediction.Price);\n});\n```\n\nThe pool manages thread-safe engine reuse for you, integrates with dependency injection, and with `watchForChanges: true` automatically reloads the model when the file changes — so you can deploy a retrained model without restarting the app.\n\n`Transform`\n\n```\nIDataView batch = mlContext.Data.LoadFromEnumerable(newHouses);\nIDataView scored = model.Transform(batch);\nvar results = mlContext.Data.CreateEnumerable<HousePrediction>(scored, reuseRowObject: false).ToList();\n```\n\nFor scoring many rows at once (nightly jobs, bulk imports), `Transform` over an `IDataView` is much more efficient than looping a `PredictionEngine`.\n\n```\npublic class SentimentData\n{\n    [LoadColumn(0)] public string Text { get; set; }\n    [LoadColumn(1), ColumnName(\"Label\")] public bool Sentiment { get; set; }   // true = positive\n}\n\npublic class SentimentPrediction\n{\n    [ColumnName(\"PredictedLabel\")] public bool IsPositive { get; set; }\n    public float Probability { get; set; }\n    public float Score { get; set; }\n}\njs\nvar data = mlContext.Data.LoadFromTextFile<SentimentData>(\"reviews.tsv\", hasHeader: true);\nvar split = mlContext.Data.TrainTestSplit(data, testFraction: 0.2, seed: 1);\n\nvar pipeline = mlContext.Transforms.Text.FeaturizeText(\"Features\", nameof(SentimentData.Text))\n    .Append(mlContext.BinaryClassification.Trainers.SdcaLogisticRegression(\n        labelColumnName: \"Label\", featureColumnName: \"Features\"));\n\nvar model = pipeline.Fit(split.TrainSet);\n\nvar predictions = model.Transform(split.TestSet);\nvar metrics = mlContext.BinaryClassification.Evaluate(predictions, labelColumnName: \"Label\");\n\nConsole.WriteLine($\"Accuracy:  {metrics.Accuracy:P1}\");\nConsole.WriteLine($\"AUC:       {metrics.AreaUnderRocCurve:F3}\");\nConsole.WriteLine($\"F1:        {metrics.F1Score:F3}\");\nConsole.WriteLine($\"Precision: {metrics.PositivePrecision:F3}   Recall: {metrics.PositiveRecall:F3}\");\n\nvar engine = mlContext.Model.CreatePredictionEngine<SentimentData, SentimentPrediction>(model);\nvar result = engine.Predict(new SentimentData { Text = \"Absolutely loved it, would buy again!\" });\nConsole.WriteLine($\"{(result.IsPositive ? \"Positive\" : \"Negative\")} ({result.Probability:P0})\");\n```\n\nLogistic-regression trainers output a **probability**, not just a yes/no — you can apply your own threshold. If a false positive and a false negative cost different amounts (see precision vs. recall in the AI/ML Fundamentals guide), tune that threshold rather than accepting the default 0.5.\n\n```\npublic class TicketData\n{\n    [LoadColumn(0)] public string Area { get; set; }    // the label: \"Billing\", \"Bug\", \"Feature\", ...\n    [LoadColumn(1)] public string Title { get; set; }\n}\n\npublic class TicketPrediction\n{\n    [ColumnName(\"PredictedLabelText\")] public string Area { get; set; }\n    public float[] Score { get; set; }\n}\njs\nvar pipeline = mlContext.Transforms.Conversion.MapValueToKey(\"Label\", nameof(TicketData.Area))   // string -> key\n    .Append(mlContext.Transforms.Text.FeaturizeText(\"Features\", nameof(TicketData.Title)))\n    .Append(mlContext.MulticlassClassification.Trainers.SdcaMaximumEntropy(\"Label\", \"Features\"))\n    // map the predicted key back to readable text, in a NEW column so evaluation can still use the key\n    .Append(mlContext.Transforms.Conversion.MapKeyToValue(\"PredictedLabelText\", \"PredictedLabel\"));\n\nvar model = pipeline.Fit(split.TrainSet);\nvar metrics = mlContext.MulticlassClassification.Evaluate(model.Transform(split.TestSet), labelColumnName: \"Label\");\n\nConsole.WriteLine($\"Micro-accuracy: {metrics.MicroAccuracy:P1}\");   // overall fraction correct\nConsole.WriteLine($\"Macro-accuracy: {metrics.MacroAccuracy:P1}\");   // average per-class accuracy\nMicro- vs macro-accuracy: micro-accuracy is dominated by the most common classes;\nmacro-accuracy weights every class equally. If rare classes matter, watch MACRO —\na large gap between the two means the model is doing well mainly on the common classes.\n```\n\nThis is the pipeline built up through Sections 3–9, in one place:\n\n``` js\nvar data  = mlContext.Data.LoadFromTextFile<HouseData>(\"houses.csv\", hasHeader: true, separatorChar: ',');\nvar split = mlContext.Data.TrainTestSplit(data, testFraction: 0.2, seed: 1);\n\nvar pipeline = mlContext.Transforms.ReplaceMissingValues(nameof(HouseData.Size))\n    .Append(mlContext.Transforms.Categorical.OneHotEncoding(\"NeighborhoodEncoded\", nameof(HouseData.Neighborhood)))\n    .Append(mlContext.Transforms.Concatenate(\"Features\",\n        nameof(HouseData.Size), nameof(HouseData.Bedrooms), \"NeighborhoodEncoded\"))\n    .Append(mlContext.Transforms.NormalizeMinMax(\"Features\"))\n    .AppendCacheCheckpoint(mlContext)\n    .Append(mlContext.Regression.Trainers.Sdca(\n        labelColumnName: nameof(HouseData.Price), featureColumnName: \"Features\"));\n\nvar model   = pipeline.Fit(split.TrainSet);\nvar metrics = mlContext.Regression.Evaluate(model.Transform(split.TestSet), labelColumnName: nameof(HouseData.Price));\n\nConsole.WriteLine($\"R²: {metrics.RSquared:F3}   RMSE: {metrics.RootMeanSquaredError:F0}   MAE: {metrics.MeanAbsoluteError:F0}\");\nMAE   (mean absolute error):  the average size of the miss, in the label's own units.\n                              \"On average we're off by about $18,000.\"  Easy to explain.\nRMSE  (root mean squared):    like MAE but penalizes LARGE misses much more heavily.\n                              RMSE much bigger than MAE => a few very bad predictions.\nR²    (coefficient of determination): fraction of the variation in price the model\n                              explains. 1.0 = perfect; 0 = no better than always\n                              predicting the average; NEGATIVE = worse than that.\n```\n\nAlways judge these against the scale of the label: an RMSE of 18,000 is excellent for million-dollar homes and terrible for $50,000 ones. And as always, compare against a baseline (predicting the training-set average) to see how much the model genuinely adds.\n\n```\npublic class CustomerData\n{\n    [LoadColumn(0)] public float Age { get; set; }\n    [LoadColumn(1)] public float AnnualIncome { get; set; }\n    [LoadColumn(2)] public float SpendingScore { get; set; }\n}\n\npublic class ClusterPrediction\n{\n    [ColumnName(\"PredictedLabel\")] public uint ClusterId { get; set; }\n    [ColumnName(\"Score\")] public float[] Distances { get; set; }   // distance to each cluster centroid\n}\njs\nvar data = mlContext.Data.LoadFromTextFile<CustomerData>(\"customers.csv\", hasHeader: true, separatorChar: ',');\n\nvar pipeline = mlContext.Transforms.Concatenate(\"Features\",\n        nameof(CustomerData.Age), nameof(CustomerData.AnnualIncome), nameof(CustomerData.SpendingScore))\n    .Append(mlContext.Transforms.NormalizeMinMax(\"Features\"))       // CRITICAL for distance-based clustering\n    .Append(mlContext.Clustering.Trainers.KMeans(\"Features\", numberOfClusters: 3));\n\nvar model = pipeline.Fit(data);\n\nvar predictions = model.Transform(data);\nvar metrics = mlContext.Clustering.Evaluate(predictions, scoreColumnName: \"Score\", featureColumnName: \"Features\");\nConsole.WriteLine($\"Average distance: {metrics.AverageDistance:F3}   Davies-Bouldin: {metrics.DaviesBouldinIndex:F3}\");\n\nvar engine = mlContext.Model.CreatePredictionEngine<CustomerData, ClusterPrediction>(model);\nvar result = engine.Predict(new CustomerData { Age = 34, AnnualIncome = 72_000, SpendingScore = 61 });\nConsole.WriteLine($\"Cluster {result.ClusterId}\");\n- NORMALIZE. K-means works on distances; an unscaled feature like income\n  (tens of thousands) swamps age (tens) and effectively decides the clusters alone.\n- You must CHOOSE numberOfClusters. There's no label to tell you the \"right\" k.\n  Try several values and compare metrics: lower AverageDistance is tighter clusters\n  (but always improves as k grows), and a LOWER Davies-Bouldin index means better-\n  separated clusters. Pick where adding clusters stops helping much.\n- The metrics only say the clusters are TIGHT and SEPARATED — not that they are\n  MEANINGFUL. A person still has to look at what each cluster contains\n  (average age, income, spending) and decide if the groups make business sense.\n- Cluster IDs are arbitrary labels: \"cluster 2\" has no inherent meaning, and the\n  numbering can change between training runs.\n```\n\nML.NET's recommender uses **matrix factorization**: it learns a small vector for each user and each item such that their dot product approximates the rating the user would give. Training data is simply (user, item, rating) triples.\n\n```\npublic class MovieRating\n{\n    [LoadColumn(0)] public float userId { get; set; }\n    [LoadColumn(1)] public float movieId { get; set; }\n    [LoadColumn(2)] public float Label { get; set; }       // the rating\n}\n\npublic class MovieRatingPrediction\n{\n    public float Label { get; set; }\n    public float Score { get; set; }                       // the predicted rating\n}\nusing Microsoft.ML.Trainers;     // for MatrixFactorizationTrainer.Options\n\nvar data  = mlContext.Data.LoadFromTextFile<MovieRating>(\"ratings.csv\", hasHeader: true, separatorChar: ',');\nvar split = mlContext.Data.TrainTestSplit(data, testFraction: 0.2, seed: 1);\n\nvar options = new MatrixFactorizationTrainer.Options\n{\n    MatrixColumnIndexColumnName = \"userIdEncoded\",\n    MatrixRowIndexColumnName    = \"movieIdEncoded\",\n    LabelColumnName             = \"Label\",\n    NumberOfIterations          = 20,\n    ApproximationRank           = 100\n};\n\nvar pipeline = mlContext.Transforms.Conversion.MapValueToKey(\"userIdEncoded\", \"userId\")\n    .Append(mlContext.Transforms.Conversion.MapValueToKey(\"movieIdEncoded\", \"movieId\"))\n    .Append(mlContext.Recommendation().Trainers.MatrixFactorization(options));\n\nvar model = pipeline.Fit(split.TrainSet);\n\nvar metrics = mlContext.Regression.Evaluate(model.Transform(split.TestSet), labelColumnName: \"Label\", scoreColumnName: \"Score\");\nConsole.WriteLine($\"RMSE: {metrics.RootMeanSquaredError:F3}\");\n\n// Predict: how would user 6 rate movie 10?\nvar engine = mlContext.Model.CreatePredictionEngine<MovieRating, MovieRatingPrediction>(model);\nConsole.WriteLine($\"Predicted rating: {engine.Predict(new MovieRating { userId = 6, movieId = 10 }).Score:F1}\");\nThe model scores ONE (user, item) pair at a time. To produce a \"top 10 for user 6\"\nlist, score that user against EVERY candidate item (usually excluding items they've\nalready rated) and sort descending by Score.\n- COLD START: a brand-new user or item has no ratings, so the model has nothing to\n  learn from — you need a fallback (popular items, content-based rules).\n- Matrix factorization learns only from IDs and ratings; it doesn't use item\n  attributes or text. (ML.NET also offers a field-aware factorization machine\n  trainer for combining extra features.)\n- Evaluate with RMSE on held-out ratings, but remember: low RMSE doesn't guarantee\n  GOOD recommendations — a ranking quality metric and, ideally, online testing\n  matter more for a real product.\n```\n\nML.NET supports two quite different styles of anomaly detection, and choosing the right one matters.\n\nBest for metrics over time: sales, latency, sensor readings, error counts.\n\n```\nusing Microsoft.ML.Data;\n\npublic class MetricPoint\n{\n    public float Value { get; set; }\n}\n\npublic class SpikePrediction\n{\n    [VectorType(3)] public double[] Prediction { get; set; }   // [alert, score, p-value]\n}\njs\nvar points = LoadMetricPoints();                                   // List<MetricPoint>, in time order\nIDataView dataView = mlContext.Data.LoadFromEnumerable(points);\n\nvar pipeline = mlContext.Transforms.DetectIidSpike(\n    outputColumnName: nameof(SpikePrediction.Prediction),\n    inputColumnName:  nameof(MetricPoint.Value),\n    confidence: 95.0,\n    pvalueHistoryLength: points.Count / 4);\n\n// These detectors learn as they stream; fit on an EMPTY dataset to create the transformer\nITransformer model = pipeline.Fit(mlContext.Data.LoadFromEnumerable(new List<MetricPoint>()));\n\nIDataView transformed = model.Transform(dataView);\nvar results = mlContext.Data.CreateEnumerable<SpikePrediction>(transformed, reuseRowObject: false).ToList();\n\nfor (int i = 0; i < results.Count; i++)\n    if (results[i].Prediction[0] == 1)                             // index 0: 1 = anomaly alert\n        Console.WriteLine($\"Spike at point {i}: value={points[i].Value}, p-value={results[i].Prediction[2]:F4}\");\nThe 3-element output vector is:  [0] alert (1 = anomaly)   [1] score   [2] p-value\n- DetectIidSpike     : spikes in data assumed independent and identically distributed\n- DetectIidChangePoint / DetectChangePointBySsa : a persistent SHIFT in level, not a one-off spike\n- DetectSpikeBySsa   : spikes in data with SEASONALITY (daily/weekly cycles), via\n                       singular spectrum analysis — needs window-size settings\nHigher confidence => fewer alerts but more are real; lower => more alerts, more noise.\n```\n\nBest for records described by several numeric features (transactions, device telemetry) with no time dimension.\n\n``` js\nvar pipeline = mlContext.Transforms.Concatenate(\"Features\", \"Amount\", \"Hour\", \"ItemCount\")\n    .Append(mlContext.Transforms.NormalizeMinMax(\"Features\"))\n    .Append(mlContext.AnomalyDetection.Trainers.RandomizedPca(featureColumnName: \"Features\", rank: 2));\n\nvar model = pipeline.Fit(normalTrainingData);                      // train primarily on NORMAL data\nvar predictions = model.Transform(newData);\n// output columns: PredictedLabel (bool: true = anomaly) and Score (higher = more anomalous)\nPCA learns what \"normal\" looks like (the main directions of variation). Rows that\nreconstruct poorly from those directions — that sit far from the normal pattern —\nscore as anomalies. Keep rank smaller than the number of features.\n- \"Anomalous\" means STATISTICALLY UNUSUAL, not WRONG. A Black Friday sales spike is\n  an anomaly and is perfectly legitimate. Anomaly detectors surface candidates for\n  a human or a rule to judge.\n- Alert fatigue is the real failure mode: too many false alarms and people ignore\n  all of them. Tune confidence/thresholds against how many alerts a team can handle.\n- Evaluating is hard because true anomalies are rare; if you have even a small\n  labeled set, use mlContext.AnomalyDetection.Evaluate with it.\n```\n\nAutoML automates the repetitive part of model selection: it runs many combinations of preprocessing and trainers against your data and reports the best one.\n\n``` js\nusing Microsoft.ML.AutoML;       // package: Microsoft.ML.AutoML\n\nvar experiment = mlContext.Auto().CreateRegressionExperiment(maxExperimentTimeInSeconds: 60);\nvar result = experiment.Execute(split.TrainSet, labelColumnName: nameof(HouseData.Price));\n\nRunDetail<RegressionMetrics> best = result.BestRun;\nConsole.WriteLine($\"Best trainer: {best.TrainerName}\");\nConsole.WriteLine($\"Validation R²: {best.ValidationMetrics.RSquared:F3}\");\n\n// Always confirm on the untouched test set\nvar testMetrics = mlContext.Regression.Evaluate(best.Model.Transform(split.TestSet), labelColumnName: nameof(HouseData.Price));\nAutoML experiments exist for: regression, binary classification, multiclass\nclassification, recommendation, and ranking — created via\nmlContext.Auto().Create<Task>Experiment(...).\n- Model Builder (a Visual Studio extension): a GUI wizard — pick a scenario, point to\n  your data, train, and it generates the C# consumption code.\n- ML.NET CLI (the `mlnet` tool): run AutoML from the command line and generate\n  a trained model and project.\nGreat for:\n  - Getting a strong baseline quickly, and learning which trainer families suit your data\n  - Teams with .NET skills but limited ML experience\n\nDoesn't replace:\n  - Understanding your data. AutoML optimizes the metric you give it; it can't tell you\n    that you have data leakage, a mislabeled column, or the WRONG metric for your business.\n  - Feature engineering and data cleaning, which usually matter more than trainer choice.\n  - A final check on a held-out test set — the validation score AutoML reports was used\n    to pick the winner, so it's slightly optimistic.\n```\n\nAutoML's API surface has evolved across ML.NET versions (including newer, more customizable experiment APIs alongside the one shown here), so check the official documentation for the version you install.\n\n| Pitfall | Why it hurts | Better approach | \n|---|---|---|\n| Sharing one `PredictionEngine` across threads / requests | It isn't thread-safe; concurrent use gives race conditions and wrong predictions | Use `PredictionEnginePool` in ASP.NET Core (Section 11) | \n| Creating a new `PredictionEngine` per request | Building one is relatively expensive and wastes performance | Pool and reuse engines via `PredictionEnginePool` (Section 11) | \n| Evaluating on the training data | The metric reflects memorization, not real-world performance | Split first; evaluate on the held-out test set (Section 9) | \n| Fitting preprocessing on the whole dataset before splitting | Test-set statistics leak into training and inflate scores | Split first, then `Fit` the pipeline on the training set only (Section 7) | \n| Column-name mismatches (\"Schema mismatch\" errors) | Trainers look for `Label` /`Features` by default; renamed columns aren't found | Pass `labelColumnName` /`featureColumnName` explicitly, or use`[ColumnName]` (Section 8) | \n| Using `double` instead of`float` for numeric columns | ML.NET's native numeric type is `float` ; mismatches cause schema errors or needless conversions | Use `float` in input/output classes (Section 3) | \n| Forgetting to normalize for linear and distance-based models | Large-scale features dominate; clustering results are driven by one feature | Add `NormalizeMinMax` /`NormalizeMeanVariance` before the trainer (Sections 5, 14) | \n| Collecting `CreateEnumerable` results with`reuseRowObject: true` | The same object is overwritten each iteration, so the list holds duplicates of the last row | Pass `reuseRowObject: false` when storing results (Section 4) | \n| Skipping `AppendCacheCheckpoint` on expensive pipelines | Lazy `IDataView` recomputes transforms on every trainer pass | Cache after costly preprocessing when the data fits in memory (Section 7) | \n| Using `CustomMapping` and then failing to load the model elsewhere | The saved model references code that isn't available in the consuming process | Prefer built-in transforms, or precompute derived features before ML.NET (Section 6) | \n| Trusting accuracy on imbalanced classification | A majority-class guesser scores high while finding nothing | Check precision, recall, F1, and AUC (Section 9, 12) | \n| Treating cluster output as automatically meaningful | Tight, separated clusters may still be business-meaningless; IDs are arbitrary | Inspect each cluster's contents and validate with domain knowledge (Section 14) | \n| Acting on every anomaly flag | \"Unusual\" isn't \"wrong\"; too many alerts cause alert fatigue | Tune confidence, route flags to review, and add business rules (Section 16) | \n| Ignoring cold start in recommenders | New users/items have no ratings to learn from | Provide a popularity or content-based fallback (Section 15) | \n| Treating AutoML's validation score as the final answer | It was used to select the winner, so it's optimistic | Confirm on an untouched test set (Section 17) | \n| Never retraining the deployed model | Inference doesn't learn; the model goes stale as data drifts | Monitor production quality and retrain on a schedule (Section 10) | \n\n| Concept | API | Purpose | \n|---|---|---|\n| Entry point | `new MLContext(seed: 0)` | Factory for all ML.NET operations; seed for reproducibility | \n| Load a file | `mlContext.Data.LoadFromTextFile<T>(path, hasHeader, separatorChar)` | Read CSV/TSV into an `IDataView` | \n| Load in-memory | `mlContext.Data.LoadFromEnumerable(list)` | Wrap a C# collection as an `IDataView` | \n| Split data | `mlContext.Data.TrainTestSplit(data, testFraction)` | Separate training and test sets | \n| Inspect data | `data.Preview(maxRows)` | Peek at schema and rows | \n| Handle missing | `Transforms.ReplaceMissingValues(col)` | Fill in missing numeric values | \n| Encode categories | `Transforms.Categorical.OneHotEncoding(out, in)` | Turn text categories into numbers | \n| Scale features | `Transforms.NormalizeMinMax(\"Features\")` | Put features on a comparable range | \n| Build features | `Transforms.Concatenate(\"Features\", cols...)` | Combine columns into one feature vector | \n| Featurize text | `Transforms.Text.FeaturizeText(\"Features\", \"Text\")` | Text to numeric n-gram features | \n| Chain steps | `estimator.Append(next)` | Compose a pipeline | \n| Cache | `.AppendCacheCheckpoint(mlContext)` | Avoid recomputing upstream transforms | \n| Train | `pipeline.Fit(trainData)` | Produce the trained `ITransformer` | \n| Apply model | `model.Transform(data)` | Score a whole `IDataView` | \n| Evaluate | `mlContext.<Task>.Evaluate(predictions, ...)` | Compute task-specific metrics | \n| Cross-validate | `mlContext.<Task>.CrossValidate(data, pipeline, folds)` | More reliable metric estimate | \n| Save / load | `mlContext.Model.Save(...)` /`.Load(...)` | Persist the whole pipeline as a `.zip` | \n| Single prediction | `mlContext.Model.CreatePredictionEngine<TIn,TOut>(model)` | One-off prediction (not thread-safe) | \n| Web serving | `AddPredictionEnginePool<TIn,TOut>().FromFile(...)` | Thread-safe, DI-friendly prediction in ASP.NET Core | \n| Binary classifier | `BinaryClassification.Trainers.SdcaLogisticRegression(...)` | Yes/no predictions with probability | \n| Multiclass | `MulticlassClassification.Trainers.SdcaMaximumEntropy(...)` | One-of-many categories | \n| Regression | `Regression.Trainers.Sdca(...)` /`FastTree(...)` | Predict a number | \n| Clustering | `Clustering.Trainers.KMeans(\"Features\", k)` | Group unlabeled items | \n| Recommendation | `Recommendation().Trainers.MatrixFactorization(options)` | Predict user–item ratings | \n| Spike detection | `Transforms.DetectIidSpike(...)` | Flag unusual points in a time series | \n| PCA anomalies | `AnomalyDetection.Trainers.RandomizedPca(...)` | Flag unusual rows in tabular data | \n| AutoML | `mlContext.Auto().Create<Task>Experiment(...)` | Automatically search trainers and settings | \n\nML.NET's design rests on a small number of ideas that, once clear, make every task feel familiar: an `MLContext` that is the root of everything, an `IDataView` that is the lazy, schema-aware table flowing through the system, and a pipeline of estimators that you `Fit` once to get a single transformer carrying preprocessing and model together. Because the whole pipeline is one saveable artifact, the transformations applied in production are guaranteed to match training — which removes one of the most common and hardest-to-spot sources of ML bugs. And because swapping a trainer is a one-line change, the five task types in this guide — classification, regression, clustering, recommendation, and anomaly detection — share one workflow rather than being five separate skills.\n\nThe framework's value is greatest when you're honest about where it fits: classic ML on tabular and text data, running inside .NET services, with models that load like any other dependency. The rest is discipline that no library can supply for you — splitting data before fitting anything, evaluating on data the model hasn't seen, choosing metrics that reflect what mistakes actually cost, treating cluster and anomaly outputs as candidates for human judgment rather than answers, and retraining as the world changes. Combine that discipline with ML.NET's consistent API, `PredictionEnginePool` for safe serving, and AutoML for a fast baseline, and you can take a model from a CSV file to a production endpoint without ever leaving C#.\n\n*Found this useful? Feel free to star the repo, open an issue with corrections, or share the \"one shared PredictionEngine gave random answers under load\" story that made the case for PredictionEnginePool click better than any documentation page.*", "url": "https://wpnews.pro/news/ml-net", "canonical_source": "https://dev.to/rhuturaj_takle/mlnet-5h4o", "published_at": "2026-10-09 17:17:18+00:00", "updated_at": "2026-10-09 17:21:21.103416+00:00", "lang": "en", "topics": ["machine-learning", "developer-tools", "ai-tools", "mlops"], "entities": ["ML.NET", "Microsoft", ".NET", "C#", "F#", "ONNX", "ONNX Runtime", "PyTorch"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ml-net", "markdown": "https://wpnews.pro/news/ml-net.md", "text": "https://wpnews.pro/news/ml-net.txt", "jsonld": "https://wpnews.pro/news/ml-net.jsonld"}}