How to Run Local AI Models on Your Own Computer

34.7K views
•
December 22, 2025
by
David Ondrej
YouTube video player
How to Run Local AI Models on Your Own Computer

TL;DR

Local AI models run entirely on your own machine (phone, laptop, or PC) using its CPU or GPU, so your prompts and data never leave your computer, every query is free, and the models work offline. The gap between frontier cloud models and open-source local models is shrinking, making now a strong time to start. Tools like LM Studio let you download and run them.

Transcript

My name is David Andre and here is how to install and run local AI models. Now we're witnessing the local revolution. Since AI is improving so fast, it is now possible to run very good models locally on your computer. Yet most people still aren't using any local models and just default to using CI GBD. So in this video you'll learn what local model... Read More

Key Insights

  • Local AI models are any AI models you can run on your own machine, from a phone or laptop to a PC. Anything with a CPU or GPU can run one, and more powerful hardware allows larger, more capable models.
  • Data privacy is a core advantage because when models run locally, all prompts, context, and attached files stay on your computer instead of being sent to a remote server, making local models ideal for sensitive or proprietary work.
  • Running local models is free at $0 per query, with no API fees, no token limits, and no monthly subscriptions, so it does not matter whether you send one prompt or one million prompts.
  • Local models work offline, so they keep functioning even when internet goes down. David uses multiple locally downloaded models on his MacBook for coding, brainstorming, and learning concepts while flying.
  • The performance gap between cutting-edge frontier models and open-source models you can run locally is shrinking, meaning local models are now roughly as good as the cutting-edge models were about a year ago.
  • Local models can be fine-tuned, uncensored, and kept free of hidden bias because they are open weight, giving you access to the parameters so you can train them further on specific data for tasks like writing in your style or building a company chatbot.
  • Artificial Analysis (artificialanalysis.ai) is an independent benchmarking platform that ranks open-source models by speed, price, output, latency, and intelligence, letting you pick the best model for the category your machine can run.
  • LM Studio is a free (unsponsored) tool for downloading and interacting with local models like GPT, Qwen, Gemma, and DeepSeek, and it offers user, power user, and developer modes for different levels of control.

Install to Summarize YouTube Videos and Get Transcripts

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What are local AI models?

Local AI models are any AI model you can run on your own machine, where machine means any type of computer, including your phone, laptop, PC, or any device with a CPU or GPU. The more powerful your computer, the more models and the more powerful models you can run locally. This is fundamentally different from cloud-based AI, where the model runs on a remote supercomputer rather than your own hardware.

Q: Why should I use local AI models instead of cloud services?

There are three main reasons. First, privacy: all your prompts, data, context, and attached files stay on your machine and never leave it, making local models ideal for sensitive or proprietary work. Second, cost: local models are free, with no API fees, token limits, or monthly subscriptions, whether you send one prompt or a million. Third, offline access: local models keep working even when your internet goes down, which is useful when flying or in any worst-case scenario.

Q: How much does it cost to run AI models locally?

Running AI models locally costs $0. Once you set up a local model, every single query and prompt you send is completely free. It does not matter whether you send one prompt or one million prompts. There are no API fees, no token limits, and no monthly subscriptions. This contrasts with cloud AI subscriptions that can add up quickly, especially higher tiers around $100 or $200 a month, and become even more expensive if you use multiple models like Claude, Grok, Gemini, and GPT.

Q: How do I find the best local AI model to run?

Go to the website artificialanalysis.ai, an independent benchmarking platform that compares AI models across speed, price, output, latency, and more. Click on models at the top, select open-source models, and you can see all the different open-source models and how they rank across metrics like intelligence and specific benchmarks. Because a new powerful open-source model is released almost every week, the best choice changes constantly, so this resource lets you check the category your machine can run and see what is currently best.

Q: What tool can I use to run local AI models?

One of the best options is LM Studio, which the creator recommends as a great, unsponsored tool. Its interface has all the same features as ChatGPT, and it lets you choose between models like GPT, Qwen, Gemma, DeepSeek, and many more. You download it from the homepage, install it by double-clicking and dragging it into your applications folder on Mac (the Windows setup is similar), and it offers three modes: user, power user, and developer, with developer giving the most options.

Q: Why aren't local AI models more mainstream?

According to the video, big tech companies and the main AI research labs don't want you to know about running AI locally. Edge compute, meaning running AI models locally on your own device, is the single biggest risk to all the hyperscalers. Their business model relies on renting you their intelligence through the cloud, so if you have a powerful enough AI running for free on your laptop, the $20-a-month subscription market completely disappears.

Q: Can local AI models be customized or fine-tuned?

Yes. Local models can be fine-tuned, uncensored, and kept on your computer where nobody can remove them or change the system prompt without you knowing. Fine-tuned models are AI models trained further on specific data to excel at a particular task, such as writing in your style, acting as your legal assistant, or serving as a custom company chatbot built on company data. Because most local models are open weight, you have access to the model weights and parameters, which is what allows you to change them and fine-tune the model.

Q: What model does David use, and what is its architecture?

David uses Nemotron 3 Nano 30B A3B, part of Nvidia's open-source Nemotron 3 family. He describes it as the only open-source local model with a 1 million context window and agentic capabilities that requires only 24GB of VRAM. It uses a hybrid Mamba-transformer architecture with a mixture of experts, meaning the model is split into specialized experts (math, language, creativity, programming) and only activates the most relevant few per query, with just 3 billion parameters active at any point. He calls it the most powerful small open-source model out right now.

Summary & Key Takeaways

  • The local AI revolution is here because open-source models have improved fast enough to run capably on personal computers. Most people still default to cloud AI, but local models are any AI you can run on your own machine, and more powerful hardware unlocks larger, stronger models.

  • The main benefits of local models are privacy, cost, and offline access: your prompts and files never leave your computer, every query is free with no subscriptions or token limits, and the models keep working without internet. They can also be fine-tuned and kept free of hidden ideological bias.

  • To start, check artificialanalysis.ai to find the best open-source model your machine can run, then use a tool like LM Studio to download and interact with it. David demonstrates Nvidia's Nemotron 3 Nano 30B A3B, a hybrid Mamba-transformer mixture-of-experts model needing only 24GB VRAM.


Read in Other Languages (beta)

Share This Summary 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator

Explore More Summaries from David Ondrej 📚

Summarize YouTube Videos and Get Video Transcripts with 1-Click

Download browser extensions on:

Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator