Accuracy
Percentage of correct predictions made by the model.
Plain-English definitions of common artificial intelligence, machine learning, model, agent, and developer terms.
Percentage of correct predictions made by the model.
Refers to AI systems that act autonomously β planning steps, calling tools, and executing multi-step tasks without human intervention at each step. An "agentic" app gives the model agency beyond single-turn Q&A.
A method or series of instructions used to generate an ML model (e.g., linear regression, decision trees, neural networks).
Application Programming Interface. A defined way for one piece of software to talk to another. In AI, an API lets your app call a model (e.g., OpenAI, DeepSeek, Groq) over HTTP and get back generated text or other output.
A secret token that authenticates your requests to a service's API. Treat it like a password: never commit it, never share it.
Browser automation CLI designed for AI agents (e.g., Vercel Labs' agent-browser). Lets an agent control a real browser β navigate pages, click elements, fill forms, extract data β via simple CLI commands. Native Rust, compact text output minimizes token usage. 35,000+ GitHub stars.
A service that gives AI agents their own email inboxes. An agent with an inbox can receive OTP codes, sign up for services, and communicate with the outside world β email becomes the agent's identity and audit trail. (agentmail.to)
The AI company behind Claude (LLM) and MCP (Model Context Protocol). Focused on "constitutional AI" and safety.
Not a widely recognized standard AI term. Could refer to a specific project, tool, or concept. (Verify.)
A quality describing a data observation (e.g., color, size, weight).
A plain-text file containing a sequence of commands for an operating system to run automatically (.bat on Windows, shell script on Linux/macOS).
The base value a model uses when all features are zero. Also: systematic error in predictions (low bias = consistent, high bias = underfitting).
A standardized test or task used to measure and compare model performance (e.g., MMLU, GSM8K). "Benchmarks" are the scores a model achieves on these tests.
A program that converses with users in natural language, typically built on top of an LLM.
OpenAI's conversational AI product, built on GPT models. The most widely used LLM interface, available via web, mobile, and API.
Anthropic's family of LLMs. Known for strong reasoning, coding, and long-context handling. Available via claude.ai and API.
Anthropic's terminal-based AI coding agent. Runs locally in your terminal, reads/writes files, executes commands, and iterates on code β similar in spirit to Pi and OpenCode but from Anthropic.
Unsupervised grouping of data into related buckets.
Command Line Interface. A text-based way to interact with a program (as opposed to a GUI). Examples: git, docker, ollama.
The process of compressing a conversation history to fit within an LLM's context window. Old messages are summarized or truncated so the most recent context is preserved. Hermes uses compaction automatically when conversations grow long.
A node-based, visual interface for running and chaining generative AI models (Stable Diffusion, Wan, etc.). You build pipelines of nodes (model loader β sampler β decoder β save) and execute them.
A table showing a classifier's performance: true positives, true negatives, false positives, false negatives.
Cascading Style Sheets. The language that controls the visual presentation of HTML (colors, layout, fonts, responsiveness). Used alongside HTML and JavaScript to build web pages and AI dashboards.
The maximum number of tokens (words/fragments) a model can process in a single conversation or request. Exceeding it means the model truncates or rejects the input.
The training state where loss barely changes between iterations.
Central Processing Unit. The general-purpose processor in a computer. Can run AI models but is much slower than a GPU for the parallel math that deep learning requires.
NVIDIA's parallel computing platform and programming model. It lets developers use the GPU for general-purpose computation β the foundation of most AI training and inference frameworks (PyTorch, TensorFlow, vLLM).
A cross-platform package and environment manager. Manages Python (and other language) environments and their dependencies. Alternative to virtualenv/pip; especially common in data science and ML.
A library or package that your project needs to function. In Python, managed via requirements.txt or pip; in Node.js, via package.json and npm.
OpenAI's always-on AI agents (launched DevDay 2026, powered by GPT-6 Astra). Unlike single-turn chatbots, Dots run 24/7, connect to 4,000+ apps, learn your preferences over time, and autonomously complete multi-step tasks. You can assign projects and they work on them in their own cloud computer.
A family of LLMs developed by the Chinese company DeepSeek (deepseek.com). Known for strong reasoning and coding performance. Served via their API or self-hosted.
A subfield of ML using multi-layer neural networks to learn hierarchical representations from data (images, text, speech).
A platform for packaging software into isolated containers (see below). Lets you bundle an app + its dependencies into a portable unit that runs the same on any machine with Docker.
A generative model (used for images, video, audio) that works by learning to gradually denoise random noise into a coherent output. The reverse of a process that slowly adds noise to data.
A folder in a file system that contains files and other directories.
A fixed-length numerical representation of text (or other data) that preserves semantic meaning. "Dog" and "animal" sit close together in embedding space.
One full pass through the entire training dataset.
An isolated software environment with its own dependencies and configuration. In Python: virtualenv/conda envs. On a server: .env files store secrets (API keys, DB passwords) outside of code.
A larger, more capable LLM intended for business/production use, often with enterprise features (security, compliance, private deployment, higher rate limits). Contrasted with smaller or open models.
A specific attribute + value in a dataset (e.g., "color is blue").
Adapting a pre-trained model to a specific task by training it on a smaller, task-specific dataset. Saves time vs. training from scratch and reduces overfitting risk.
Two competing networks: a generator creates fake data, a discriminator tries to tell real from fake; they improve until the discriminator can no longer distinguish them.
AI that creates novel content (text, images, video, code) by learning patterns from large training data rather than following predefined rules.
A family of transformer-based language models trained on massive datasets to generate human-like text. The architecture (transformer + autoregressive generation) underpins most modern LLMs.
A web platform for Git-based version control. Hosts repos (repositories), code review, issue tracking, CI/CD (Actions), and package registries.
A company that builds dedicated AI inference hardware (LPU β Language Processing Units) optimized for fast LLM token generation. Their API offers very high tokens-per-second for supported models.
xAI's (Elon Musk's company) LLM and chatbot, available via the x.com app and API. Known for real-time information and a distinctive voice.
xAI's agentic AI bots built on top of Grok. They go beyond single-turn Q&A to execute multi-step tasks, browse the web, and operate tools autonomously on the user's behalf.
General Language Model. A family of LLMs by Zhipu AI (China), known for strong reasoning and multimodal capabilities. Also a coding agent (GLM-4).
A programming language developed by Google. Known for simplicity, concurrency, and fast compilation. Used for building servers, CLIs, and system tools.
Google's family of LLMs. Multimodal (text, image, audio, video). Available via the Gemini app and Google AI Studio API.
Google's AI notebook/research tool (formerly NotebookLM). Upload documents, ask questions, generate summaries β grounded in your uploaded sources.
Gemini's autonomous research feature. You give it a research question and it browses the web, synthesizes sources, and produces a detailed report.
Graphical User Interface. A visual, click-based interface (as opposed to CLI). Examples: a web dashboard, a desktop app, ComfyUI's node editor.
When an LLM produces a plausible-sounding but factually incorrect answer instead of saying "I don't know."
A wrapper or test framework that drives a model: feeds it prompts, collects outputs, and scores results. In evaluation, an "eval harness" (like lm-eval-harness) standardizes how benchmarks are run and compared. "DeepSeek harness" likely refers to a test harness specifically built for evaluating DeepSeek models.
An AI agent framework (the one you're using right now). It connects LLMs to tools, memory, messaging platforms, and automation β enabling agentic workflows beyond simple chat.
HyperText Markup Language. The standard markup language for structuring content on the web. Used for building web pages and, in AI context, for rendering chatbot responses, dashboards, and generated content.
The process of using a trained model to generate predictions or outputs (as opposed to training, where the model learns). Running a model on your GPU to answer questions is inference.
Apple's mobile operating system for iPhone and iPad. Runs AI tools like the ChatGPT and Gemini apps; also a target platform for AI apps.
A scripting language that runs in browsers and on servers (via Node.js). Used for web interactivity, building web UIs, and server-side logic.
JavaScript Object Notation. A lightweight, human-readable data format for transmitting structured data. The standard format for API requests/responses.
Node Package Manager. The default package manager for Node.js/JavaScript. Installs dependencies from the npm registry. (npm install, npm run, etc.)
The "answer" portion of an observation in supervised learning (e.g., "cat" vs. "dog").
The size of each update step during training. Too high: overshoots; too low: trains very slowly.
A neural network trained on massive text corpora to understand and generate natural language. Examples: GPT, Claude, Gemini, DeepSeek, Qwen.
A desktop application (Windows/macOS/Linux) for running local LLMs. Provides a GUI for downloading models, configuring inference, and chatting β no code required.
The sum of errors between predicted and true values. Lower loss = better model (barring overfitting).
A field where algorithms learn from data and improve performance on tasks without being explicitly programmed.
An open protocol by Anthropic that standardizes how AI models connect to external tools, data sources, and systems. An MCP server exposes tools/resources to a model; the model calls them like functions. Replaces ad-hoc integrations with a uniform interface.
A family of efficient, high-performance LLMs by the French company Mistral AI. Known for strong performance-to-size ratio. Available via API or self-hosted. (Note: "Mistral" not "Mistrial.")
A data structure storing learned weights and biases from training. In LLM context, "model" also refers to the model name (e.g., "Qwen3.8-27B") and its file (e.g., a GGUF).
A document that accompanies a published ML model. Describes its intended use, training data, performance, limitations, ethical considerations, and usage guidelines.
The process of teaching a model on data: feeding it examples, adjusting parameters, and measuring performance until it learns the pattern. The counterpart to inference (using the trained model).
The branch of AI focused on processing, understanding, and generating human language (translation, sentiment analysis, speech recognition, etc.).
A structure of interconnected layers (neurons) modeled after the brain, used to recognize patterns in data.
The leading maker of GPUs, which are the primary hardware for AI training and inference. Their CUDA ecosystem dominates AI computing.
Baseline accuracy from always predicting the most frequent class.
A markdown-based note-taking and knowledge management app. Notes are plain .md files stored in a "vault." Supports bidirectional links, plugins, and graph views. (You're using it right now.)
An open-source, self-hosted web interface for LLMs. Provides a ChatGPT-like UI that can connect to various backends (Ollama, OpenAI-compatible APIs, etc.).
The company behind ChatGPT, GPT models, DALL-E, and the API. One of the most prominent AI companies.
A tool for running LLMs locally on your machine. Makes it easy to download, serve, and chat with models via a local API. CLI-driven; pairs well with Open WebUI or LM Studio for a GUI.
An LLM whose weights are publicly released so anyone can download, use, modify, and redistribute it (subject to its license). Examples: Llama, Mistral, Qwen, DeepSeek. Contrasted with closed/enterprise models.
Model learns training data too well, including noise; performs great on training data but poorly on new/test data.
An observation that deviates significantly from the rest of the dataset.
An open-source agent harness (the "exoskeleton" around an LLM). It wraps a model with memory, tools, triggers, instructions, and output channels so the model can operate as an autonomous software agent rather than a one-turn assistant. Microsoft runs it natively on Windows. (docs.openclaw.ai)
An open-source AI coding agent/CLI that runs locally. Similar in spirit to Claude Code but open.
A learned value (weights, coefficients) in a model. "Model size" in AI is expressed in parameters: a "7B" model has 7 billion parameters, a "27B" model has 27 billion. More parameters generally = more capability but more memory/compute required.
An open-source, terminal-based AI coding agent and agent harness by Earendil Works (Mario Zechner). Small extensible core with four default tools (read, write, edit, bash); customizable via TypeScript extensions, skills, and packages. Connects to 15+ model providers. Runs locally on your machine. (pi.dev)
When the model predicts "positive," how often is it correct? TP / (TP + FP).
The input text (instructions, question, context) you give to an LLM to guide its output.
Crafting effective input instructions for an LLM to get the best possible output. Requires understanding how LLMs work, their training data, and their limitations.
The most popular programming language for AI/ML. Dominates data science, machine learning, and scripting. Libraries: PyTorch, TensorFlow, scikit-learn, pandas, numpy.
The process of reducing the numerical precision of a model's weights to make it smaller and faster. A 16-bit (FP16) model quantized to 4-bit (e.g., GGUF Q4_K_M) uses ~4x less memory with some accuracy tradeoff. Essential for running large models on consumer hardware.
A family of LLMs developed by Alibaba. Known for strong multilingual performance, coding, and reasoning. Available via API (DashScope) or self-hosted. (You're running Qwen3.8-27B right now.)
A technique where the LLM retrieves relevant information from an external knowledge base (documents, databases) before generating a response. Reduces hallucinations by grounding answers in retrieved context.
Random Access Memory. The computer's fast, volatile working memory. In AI, models and data must fit in RAM (system memory) when not on GPU. Insufficient RAM forces slow disk swapping.
For all true positive cases, how many did the model catch? TP / (TP + FN).
Predicting a continuous numerical value (price, temperature) rather than a category.
Repository. A Git-based project directory containing code, history, branches, and metadata. Hosted on GitHub/GitLab/Bitbucket.
An isolated environment where code or processes run without affecting the host system. In AI, a sandbox protects against dangerous code execution by LLM agents.
A reusable, self-contained unit of capability for an AI agent: a set of instructions, tools, and knowledge for a specific task type (e.g., "obsidian," "weather"). Hermes uses skills to extend agent abilities.
When the true label is negative, how often is the prediction correct? TN / (TN + FP).
Training on labeled data (inputs paired with correct outputs).
A program that provides a text interface to an operating system. Synonymous with CLI/Shell. Where you type commands.
A messaging app/platform. In AI context, it's a common interface for chatting with AI agents (like this one). Hermes can send/receive messages via Telegram.
Unseen data used to evaluate how well a model generalizes.
A piece of text that a model processes: roughly 3/4 of a word in English. Models generate and count in tokens. A 1000-word response β 1300β1400 tokens.
When an LLM agent invokes an external function/tool (e.g., search, calculator, file read) as part of generating a response. The model outputs a structured request; the system executes it and returns the result.
The process of teaching a model on data: feeding it examples, adjusting parameters, and measuring performance until it learns the pattern.
A deep neural architecture built from encoder and decoder components, designed for sequential data (language, time series). The backbone of all modern LLMs.
Model is too simple; performs poorly on both training and test data.
A memory architecture (used in Apple Silicon and some mobile chips) where GPU and CPU share the same physical RAM pool, eliminating data copying between separate memory spaces. Simplifies AI workloads on Apple devices.
Finding patterns in unlabeled data with no predefined labels (e.g., clustering).
A specific application or scenario where an AI system is deployed to solve a problem. (e.g., "RAG for a legal research assistant.")
Data checked during training to catch overfitting.
Video RAM. The dedicated memory on a GPU. In AI, VRAM is the most critical hardware constraint: the model + its context + activations must fit in VRAM. A 12GB GPU (RTX 3060) can run quantized 7Bβ14B models; a 24GB GPU (RTX 3090) can run larger models or longer contexts.
How tightly clustered the model's predictions are. High variance (with low bias) suggests overfitting.
Microsoft's desktop operating system. Runs AI tools (LM Studio, Ollama, Docker Desktop, Claude Code) with some differences from Linux (paths, shell, GPU drivers).
A family of video generation models by Alibaba (Wan 2.2 I2V, etc.). Used in ComfyUI pipelines for text-to-video and image-to-video.