Free course ยท Intermediate
Meta AI and Llama Mastery โ Open Models You Can Run Yourself
What you will learn
- Explain what open weights let you do that a hosted model does not
- Run a model on your own machine and see what it costs in memory
- Write a prompt that suits a small model rather than fighting it
- Build one tool that works with no network connection
- Measure a model on speed and memory before promising it to anybody
- Chunk a long document yourself, because you own the context window
Course curriculum
4 modules ยท 15 lessons ยท a worked example and a practice task in every lesson
Module 1Foundations โ what Llama is1 of 43 lessons
Week 1 โ meet the tool, get an account, and learn the screen before you learn the prompting.
Meet Llama โ what it is and who makes it
Llama is Meta's family of open-weight models you can download, host and fine-tune. It is at its best at being yours โ the weights can be downloaded, run on your own machine, tuned on your own data and shipped in your own product. It is at its weakest at convenience: nothing is handled for you, and a small local model is measurably worse than a frontier one โ and knowing both halves is what separates somebody who uses it well from somebody who trusts it blindly.
The lesson
Llama is a family of open-weight models you can download, host and fine-tune made by Meta. The models behind it are the Llama open-weight family, in several sizes, plus the hosted assistants Meta runs on top of them. None of that matters on its own โ what matters is that you know what kind of worker you have hired. At being yours โ the weights can be downloaded, run on your own machine, tuned on your own data and shipped in your own product is the job you hand it. At convenience: nothing is handled for you, and a small local model is measurably worse than a frontier one is the job you keep.
Most people come to a course like this expecting a list of magic words. There is no such list. What there is, is a tool that produces whatever you build around it, which is the point and also the work, and produces it at a speed no human matches โ which means the skill is not in writing the prompt, it is in knowing what a good answer looks like so you can tell the difference.
Here is the first thing to try, worded the way you will word things for the rest of the course:
Explain the difference between a hosted AI model and an open-weight model as if I am a second-year student, in five bullets, and then tell me one thing I could build with the open one that I could not build with the hosted one.Read the answer twice. The first read is for the content; the second is for the shape โ did it answer the question you asked, or the question it found easiest? That second read is the habit this whole course is built on.
Example The last clause is the useful half โ it turns a definition into a reason to care.
Practice Open Llama, ask it the one question you would normally put to a search engine about your own field, and write down two things: whether the answer was right, and whether it would have taken you longer to find it yourself.
Signing in: what is free, what is paid, and what you actually need
You do not need the paid plan to finish this course. Start on the free tier โ the weights are free to download and use under Meta's community licence, and the hosted assistants are free to chat with. Upgrade only when you hit a wall you can name: a longer file, a newer model, or a rate limit you keep meeting.
The lesson
Every one of these tools has a free tier that is good enough to learn on and a paid tier that removes a limit. there is nothing to buy from Meta โ the costs are your own hardware, or an hour of a rented GPU. The mistake is buying the paid plan in week one, before you know which limit you hit โ you end up paying to remove ceilings you were never going to touch.
Work out your own honest usage first. How many questions a day do you actually ask? How big are the files you upload? Do you need the newest model, or the fast one? For revision, for drafting, for coursework, the answer is usually the free tier.
If your college or workplace provides an account, use it: a licensed commercial deployment, or a managed provider serving the same weights is how most people in a job get access, and asking your placement cell whether one exists costs nothing.
Example A 8B model runs acceptably on a modern laptop with 16GB of memory, which means you can experiment with no cloud account at all.
Practice Create the account, find the plan page, and write down in one line which limit you would hit first in your own week. It is usually a message cap or a file-size cap, not the model.
The screen: where every control lives
A tour of the interface you will live in โ two very different surfaces: a normal chat assistant, and a command line where you load a file of weights and talk to it. Every panel has a reason to exist, and half of them are the difference between a chat and a system.
The lesson
The interface of Llama is two very different surfaces: a normal chat assistant, and a command line where you load a file of weights and talk to it. That sentence is worth slowing down on, because the single biggest cause of bad output is not a bad prompt โ it is a good prompt typed into the wrong place.
The history list is your memory of what worked. Name your conversations. The temporary or private mode is for anything you would not want in an account's history. The settings panel holds the personal instructions that apply to everything, which is where your context belongs rather than repeated at the top of every message.
Do this once, properly: run one prompt against two different model sizes on your own machine and time both, so you can feel the trade you are making. It takes fifteen minutes and saves you those fifteen minutes every week after.
Example The chat assistant is a normal assistant. The interesting surface is the terminal, where you choose the model size and see what it costs you in memory.
Practice Spend fifteen minutes doing nothing but clicking. Open every panel, rename one conversation, and save one setting you will want again. Fluency with the screen is what stops you re-explaining yourself every session.
Module 2Prompting โ getting a real answer2 of 44 lessons
Week 2 โ how Llama reads text, the four-part prompt, your own work, and what to do when the answer is wrong.
How Llama reads what you type
Context, instructions and roles. You control the context window directly โ how many tokens the model can see at once โ and you pay for it in memory, not in money. Understanding what the model can see โ and what it has already forgotten โ explains almost every disappointing answer you will get.
The lesson
You control the context window directly โ how many tokens the model can see at once โ and you pay for it in memory, not in money. This is the machinery. A model does not remember your last conversation the way a person does; it is handed text and asked to continue it well. Everything you want it to know has to be in that text, in the same window.
There is a difference between a system instruction โ the standing context, set once โ and a message, which is this request. Put your standing context in the standing place. "I am a final-year mechanical engineering student applying for data roles" belongs in your profile, not typed again at the start of every chat.
And there is a hard limit. When a conversation gets long enough, the earliest turns fall out of view. Symptoms: it contradicts an instruction you gave ten messages ago, or forgets the file you uploaded. The fix is a new conversation with a written summary of where you got to, not a longer argument.
Example Set too small a context and the model forgets the start of your own document. That limit is a number you chose, which is a very different feeling from a black box.
Practice Ask the same question twice: once on its own, once after a short paragraph of setup about who you are and what you need. Compare the two answers and write down the difference. That gap is the thing you are learning to control.
The four-part prompt: role, task, context, format
Almost every good prompt has four parts: who the model should be, what it must do, the facts it must use, and the shape of the answer you want. Leave out the fourth and you get an essay when you wanted a table.
The lesson
Write the four parts as four lines, in this order. ROLE: who it should answer as. TASK: the single thing you want done, as an instruction, not a wish. CONTEXT: the facts, pasted, not referred to. FORMAT: the exact shape of the answer โ "a table with three columns", "five bullets, no more than twelve words each".
FORMAT is the part everybody skips and the part that saves the most time. An answer you have to restructure by hand was not really an answer. Ask for the shape you are going to use: if it is going into a slide, ask for slide bullets; if it is going into a spreadsheet, ask for rows.
Here is the same request done both ways โ first as people usually write it, then in four parts:
Weak: tell me about data analyst jobs Strong: ROLE: A hiring manager for entry-level analytics roles in India. TASK: List what you screen for in a fresher's first 30 seconds. CONTEXT: I am a 2026 B.Com graduate with Excel, basic SQL and one dashboard project. No internship yet. FORMAT: A table: skill | what a fresher shows | what most get wrong. Five rows maximum. No introduction.The second prompt is not longer because long is good. It is longer because it contains four things the model cannot guess, and the guess is where the useless answer came from.
Example A smaller local model needs blunter prompts than a frontier model: short sentences, one job per request, an explicit format. Learning that makes you better at prompting everywhere.
Practice Take a task you did last week without AI and write the prompt in four labelled lines. Then run it, and rewrite only the part that failed โ not the whole prompt.
Llama for the work you actually have
Assignments, revision, email, applications, meeting notes. A private summariser that never sends a word to anybody else โ which is the honest reason a student would run a model locally at all. The test of a tool is whether it removes an hour from your week, not whether the demo looked clever.
The lesson
The highest-value use of Llama for a student or a fresher is not writing essays. It is compression: turning a 40-page chapter into the six things you actually have to remember, turning a messy set of notes into a revision sheet, turning a job description into a list of what to prove.
The second highest is structure: given a blank page, ask for three possible outlines and pick one. Being stuck is usually a problem of options, not of effort.
Here is a prompt worth keeping verbatim โ it is the one that turns a document into something you can study:
Summarise the text between the markers. Use only the text provided. Do not add information. Output exactly: SUMMARY: three sentences KEY POINTS: five bullets under 12 words UNSUPPORTED: anything you could not confirm from the text ---TEXT--- {paste} ---END---Notice the last line. Asking for the gaps is asking the tool to mark its own work, and it is the single most useful line you can add to a study prompt: the material it could not summarise is the material you have not understood yet.
Example Write the request for a local model: one job, a fixed output format, and an instruction to answer only from the supplied text.
Practice Pick the one task you repeat every week โ the one that is boring rather than hard โ and rebuild it in Llama today. Time it. Then keep the prompt that worked, saved and named.
When the answer is wrong: iterate instead of restarting
A bad answer is information. Small models fail differently: they loop, they drift off format, and they answer from imagination when the text did not say. Keep the conversation, name the fault, and correct one thing at a time โ restarting from scratch throws away everything the model has already got right.
The lesson
There are four things that usually went wrong, and each has a different fix. It answered a different question โ restate the TASK as one sentence. It made things up โ supply the facts yourself and say "use only these". It wrote too much โ ask for the length first. It sounded like a machine โ ask it to rewrite for one specific reader and cut every third word.
Say the fault out loud, in the message. "This is too long for a WhatsApp message and it sounds like a brochure." A model cannot fix a problem you have not named, and naming the problem is also how you find out what you actually wanted.
Follow-ups that work: "shorter", "only the parts that are true for a fresher", "rewrite the second sentence three ways", "what would you have to check before I send this?". That last one is worth using before anything goes out with your name on it.
Example If a local model invents an answer, lower the temperature and forbid it explicitly from adding anything. Both are settings you own here.
Practice Take the worst answer you have received this week and correct it in three follow-up messages without retyping the original prompt. Notice how much faster it converges.
Module 3Files, tools and your own data3 of 44 lessons
Week 3 โ documents and tables, Running and tuning a model on your own machine, the API, and one boring task automated.
Files, tables and long documents
Uploading a PDF, a spreadsheet or a screenshot and asking questions about it: You can feed it documents, but on a local model the practical limit is your own memory โ a document bigger than the context window has to be split by you. This is where these tools stop being a chat and start being work, and where the failure modes are worth knowing.
The lesson
You can feed it documents, but on a local model the practical limit is your own memory โ a document bigger than the context window has to be split by you. The two failures to expect are the ones nobody warns you about: a long document is summarised as it is read, so a detail on page 60 can be missed; and a table with merged cells or a scanned page is likely to be read wrongly.
So the working method is: ask narrow questions, and ask the tool to quote the line it is answering from. "From the attached file only, what does clause 4.2 let me do? Quote the sentence." If it cannot quote it, it did not read it.
For a spreadsheet, ask for the formula and an explanation instead of the computed column โ ask it to restate the question in its own words before answering, which catches most misreads โ then you can fix it yourself next month when the columns change.
Example A 200-page PDF on a small local model becomes twenty chunks and a script. That script is exactly the exercise this module is for.
Practice Upload one real document you already own โ a syllabus, a marksheet, a project report โ and ask three questions whose answers you already know. When it gets one wrong, read the passage it quoted before you blame the file.
Running and tuning a model on your own machine โ Llama's own feature
The open weights are the feature. You can download a model, run it with no network connection, point it at your own files, and โ if you get that far โ tune it on your own data so it answers in a way no hosted assistant can.
The lesson
Because it changes what you can build. A hosted assistant cannot be shipped inside your product, cannot be told to work with no internet, and cannot be tuned on a company's private data. Open weights can, and the interview sentence "I ran and fine-tuned a model" is not something most freshers can say.
Start with a small quantised model and a local runner. Get it answering at all before you optimise anything. Then measure: how much memory, how fast, how good on five questions you know the answers to. Only then decide whether a bigger model is worth the slowdown.
Here is the shape of it, in the form you will actually use:
# One-time: install a local runner (Ollama is the simplest start) # -> https://ollama.com/download ollama pull llama3.2:3b # small, fast, runs on a laptop ollama run llama3.2:3b # interactive chat in the terminal # A single question, scriptable, with no network involved: ollama run llama3.2:3b "Summarise this in three sentences: [text]" # See exactly what is on disk and how big it is: ollama listThe mistake is judging a 3B model by the standard of a frontier model and concluding that open models are bad. They are a different tool: smaller, private, free to run, and yours to change.
Example A summariser that runs offline on your laptop and never sends a word anywhere is a two-evening project, and it is the clearest demonstration of what open weights mean.
Practice Run one model locally, build one prompt around it, and write down the numbers: memory used, seconds per answer, and how it compared with a hosted model on the same job.
A local API and the hosted endpoints and your first script
A local runner exposes an OpenAI-compatible endpoint on your own machine, and the same weights are served by many managed providers, so code written for one usually moves to the other. You do not need this to finish the course, and you do not need it for a fresher job either โ but an afternoon here is what turns "I have used Llama" into "I have built with it", which is a different sentence in an interview.
The lesson
A local runner exposes an OpenAI-compatible endpoint on your own machine, and the same weights are served by many managed providers, so code written for one usually moves to the other. The idea is simple: the same model you have been chatting with also answers a web request, so you can put it inside a script, a spreadsheet or a page. A key identifies you; a request sends the text; a response comes back as data.
The reason to try it once, even if you never build anything: it makes the chat version less mysterious. You see that the whole conversation is text in and text out, that your instructions are literally lines of a request, and that "the model" is one parameter among several.
A first call looks like this:
import requests r = requests.post( "http://localhost:11434/api/generate", json={"model": "llama3.2:3b", "prompt": "Explain a database index in two sentences.", "stream": False}, ) print(r.json()["response"])The mistake is assuming local means unlimited. The model answers as fast as your hardware allows, and one long document can take minutes โ measure before you promise anybody a live feature.
Example A script that loops over your own notes folder and returns a summary file is the standard first project โ it runs offline, and it costs nothing.
Practice Get a key, run one request that works, and change one word in it to see the answer change. That is the whole of the first afternoon.
Automate one boring task with Llama
Automation is not about building a system. It is about doing one repetitive job the same way every time, in less time than last time, and being able to do it again next month.
The lesson
Pick the task by how often it happens, not by how impressive it would be. A weekly report you can half-generate beats a clever pipeline you build once and never open again.
Summarising a folder of your own notes, or turning a long text into a fixed format, entirely offline.. Whatever you choose, write the steps back out in plain English afterwards โ "Step 1, open the sheet, Step 2, paste the names โ" because the written steps are what you follow when the tool changes next quarter.
And keep a copy of the prompt next to the task. A prompt that lives only in your chat history is a prompt you will rewrite from scratch in March.
Example A script that walks your notes directory and writes one summary file per week is small, useful and finishes in an evening.
Practice Name the task you repeat most often that involves typing, then write the prompt for it and run it three weeks in a row from the same saved place. Three runs is the point at which you know whether it is genuinely automated.
Module 4Career, projects and honesty4 of 44 lessons
Week 4 โ a finished portfolio project, the privacy rules, the limits, and the Meta certification paths.
Build the portfolio project: an offline summariser for your own documents
Build a script that walks a folder, summarises every file through a local model, and writes one combined markdown file. It must run with the network switched off. It is the thing you will talk about in the interview, so it has to be small enough to finish in a fortnight and concrete enough to show a person in one minute.
The lesson
Build a script that walks a folder, summarises every file through a local model, and writes one combined markdown file. It must run with the network switched off. Finish it before you start the next one. A half-built idea shows nothing; a small finished thing shows that you can finish.
Use the tool as a collaborator, not an author: ask it for a plan, a critique and a checklist, and write the work yourself. In the interview the questions will be about the decisions โ why this, why not that โ and only the work you did yourself has answers.
Write one paragraph beside the project: what problem it solves, what you used, and what you would do differently next time. That paragraph is the interview.
Example Finished, it is a folder of documents in and one markdown summary out, plus a note of the memory and time it took.
Practice Run it twice: once with the smallest model and once with the largest that fits. Write the one paragraph comparing them โ that paragraph is the interview story.
Ethics, privacy and what never to paste
These tools send what you type to somebody else's computer and keep it in a history. Never paste passwords, government ID numbers, bank or card details, medical records, or another person's private data โ and never paste a company's confidential document.
The lesson
There is no version of this tool where your text stays on your laptop. Everything you type is sent to a server, kept in a history you can usually see, and may be reviewed or used to improve the product depending on the plan.
Check that: a local runner that silently calls a hosted endpoint is not private, and a model downloaded from a random mirror is not trustworthy. Take weights from the publisher or a well-known distribution.
The practical rule for a fresher: replace the real thing with a stand-in. "Client A", "my friend's phone number", "the amount in the offer letter". The tool rarely needs the real value to do the work, and the stand-in costs you nothing.
And the professional rule: running locally is what you do with documents you are not allowed to upload anywhere. If your employer has an approved plan or a policy, that policy is the answer, not your judgement about how sensitive a file really is.
Example This is the one course where the honest answer to "where does my data go" is nowhere, provided you actually run it locally and the model really is offline.
Practice Go through your last five conversations and delete anything containing a real phone number, a client name, or a document you did not write. If you cannot find the delete button, that is the lesson.
Limits, hallucinations and how to check
These tools predict plausible text. A small open model is weaker than a frontier model, and the gap is largest exactly where it matters most: facts, arithmetic and reasoning chains That is a mechanism, not a moral failing โ and the habit it demands is the habit of asking "where did this come from?" out loud, every time, before you use an answer.
The lesson
A model does not look things up unless it has been given a way to look things up, and even then it can attach a real number to the wrong claim. A small open model is weaker than a frontier model, and the gap is largest exactly where it matters most: facts, arithmetic and reasoning chains
So the rule is: numbers, dates, names, citations and legal or medical claims get checked in a primary source before they leave your hands. Everything else โ drafts, structure, explanations, practice โ is fair game.
Three questions to ask before you trust an answer. Where did this come from? What would make it false? Who is the original source, and can I open it? If the third one has no answer, you have writing material, not facts.
Treat a local model as an assistant for structure and rewriting, not as a source. Anything factual comes from the text you gave it or from a primary source.
Example Ask a small local model three questions you know well. It will answer all three confidently, and it will be wrong on at least one in a way you can see immediately โ which is the best possible training.
Practice Ask for one statistic with its source, then open the source. Sometimes it exists. Sometimes the citation is invented, and the number is close enough to a real one to be dangerous. Either way you will remember the exercise.
Get certified: the official Meta paths
A Way2Fresher certificate for this course is free and lives on this site. Beyond it, Meta publishes its own learning and certification material โ and the rail on this course page links to it.
The lesson
There are two things called a certificate and they are not the same. The one this site issues records that you finished a structured course and built the project at the end of it โ it is free, and it is yours to print. The ones Meta issues record that you passed their own material.
Meta publishes the model cards, licences and responsible-use guides, and distributes the weights through its own channels. There is no Meta certification for Llama; the portable credential in this area is a general cloud or machine-learning fundamentals certificate, which pairs well with a project.
Get both, in that order. The project is what an interviewer asks about; the certificate is what gets past a filter that looks for keywords. The links to the official paths are in the rail beside this lesson, each labelled with who issues it.
And put the work on the certificate, not the other way round: a certificate with no project behind it is a line on a resume, and it lasts exactly until the first technical question.
Example Explain in plain words what a context window and a quantised model are, and be able to show your own script. That explanation is worth more than a certificate in a technical interview.
Practice Finish every lesson here, take your Way2Fresher certificate, then open one official path and work through it with the project you have already built. Being certified in the tool you can already use is a small additional step.
About Meta AI and Llama Mastery โ Open Models You Can Run Yourself
Understand what open weights actually let you do, run a model on your own laptop, and build one small tool that would not have been possible through a hosted chat window.
Students who can read a little code and want to know how these models work underneath, and freshers targeting AI-adjacent roles where "I have run and tuned a model" is a real differentiator.
What you will be able to do at the end
- Explain what open weights let you do that a hosted model does not
- Run a model on your own machine and see what it costs in memory
- Write a prompt that suits a small model rather than fighting it
- Build one tool that works with no network connection
- Measure a model on speed and memory before promising it to anybody
- Chunk a long document yourself, because you own the context window
How the course is structured
4 modules and 15 lessons, arranged so each one ends with something you have built. Every lesson carries a worked example and a practice task โ the practice is the course, the reading is only the setup. Plan for 5 weeks ยท about 4 hours a week.
The full syllabus โ every lesson, its example and its practice task โ is in the Course curriculum below. Nothing is locked and nothing needs an account.
Your weekly routine
- Four sessions a week of fifty minutes: one reading, one running, one measuring, one writing up what you found.
- Always write down memory and seconds. Numbers turn an experiment into evidence.
- Keep one folder of prompts that worked on the small model โ it is a different style of writing and worth having.
What you will have built by the end
- Build a script that walks a folder, summarises every file through a local model, and writes one combined markdown file. It must run with the network switched off.
- A chunk-and-summarise pipeline for one long report
- A comparison note between a local model and a hosted one, with numbers
Where this leads for a fresher
- Machine-learning and AI engineering internships
- Backend roles that ship an AI feature inside a product
- Data roles where the data cannot leave the building
- Any technical role where "I have run and tuned a model" separates you from the queue
Titles vary between companies; the evidence does not. A deployed project, a set of queries you can explain, or a case study with real testing behind it is what a fresher interview has to work with.
Frequently asked questions
Do I need a graphics card?
Not for a small model. A modern laptop with 16GB of memory runs a 3B model acceptably. A GPU makes the larger models usable, and renting one by the hour for a few days is enough for this course.
Is a local model good enough to use daily?
Not compared with a frontier model, and that is not the point. It is good enough to be private, free and yours, which are different advantages.
How is this different from the other AI courses here?
Every other course uses somebody else's model over the internet. This one is about owning the model โ which is the only route to privacy, offline work, product integration and fine-tuning.
What will I have at the end of this course?
Three things: an offline summariser for your own documents, a saved set of prompts you wrote and tested on your own work, and a Way2Fresher certificate naming the course. Meta also publishes its own learning material for the tools it makes, and the rail on this course page links to it.
Not sure which of these you need first? The free Career Pulse check scores your skills, communication and goal clarity in about three minutes and tells you which gap to close first. Take the free check.