What ChatGPT actually does
When you ask ChatGPT or a similar AI a question, what happens inside can be said in one line: it guesses again and again what the next word should be.
That sounds cheap, but it is not. To say correctly which word is most likely after "the capital of Bangladesh", a program has to understand a good deal of geography, Bangla grammar and sentence structure. After doing this guessing over hundreds of millions of sentences, the Model ends up holding something that we, from outside, call "knowledge".
This program is called a Large Language Model, LLM for short. Large means huge in size, and Language Model means a pattern or mould of language. ChatGPT, Claude, Gemini and DeepSeek are all apps built on top of some LLM.
One thing is worth making clear right here. "ChatGPT" and "LLM" are not the same thing. The LLM is the engine, and the app is the whole car: it has a way to keep the chat history, a way to read files, and often a way to search the internet. Many of the features you see belong to the car, not the engine.
Token: to an AI, text is pieces
A Model does not see whole words. It first breaks the text into small pieces, and each piece is called a Token.
A Token is not always one word. In English "school" may be a single Token, but "unbelievable" may break into three. Punctuation, spaces and even half a word can each be a separate Token.
The rule for breaking text is fixed beforehand. The Model then turns every Token into a number, because in the end a computer understands nothing but numbers. When it writes an answer, it also makes one Token at a time, and as these join together you see sentences on the screen. This is why an AI answer arrives slowly like a typewriter: the whole thing is not made at once.
Why do you need to know about Token? Because almost everything in the AI world is measured in this unit. How much it can remember at one time is measured in Token. If you use it through an API, the bill is measured in Token. How long the answer will be is also measured in Token. If you do not understand Token, you will not understand any of these three.
A hardware shop in Khulna: the real Token arithmetic
Let us look at a real calculation so this becomes something you can hold.
The owner of a hardware shop in Khulna decides to get Bangla descriptions written for 120 products, to post on Facebook. For each product he writes an instruction of about 40 words, and the AI returns a description of about 60 words. So each product is roughly 100 words.
Now let us take a rough assumption: in Bangla one word breaks into three Token on average. This changes with the app and the Model, so think of it as a number taken for convenience, not a rule. Then each product costs about 300 Token.
For 120 products at 300 each, the total is 36,000 Token. That sounds small. But the real point is hidden right here.
If he does all 120 products in a single chat, the whole earlier conversation has to be sent again each time. At the second product the Model reads 600 Token, at the tenth 3,000, and at the 120th product it reads 36,000 Token for that one message alone. Added up, the whole chat reads about 2,178,000 Token.
And if he opens a new chat after every ten products? Each small chat reads about 16,500 Token, and twelve chats total about 198,000. That is about one eleventh.
The results are three, and he will notice all three: answers come slowly in a long chat, the API bill goes up, and after a while the AI can no longer see the instructions from the beginning.
Why Bangla takes more Token
The "three Token on average" in the calculation above is not a coincidence. The rules for breaking text into Token in most Models were built mainly by looking at English text. So a common English word is often a single Token, but a Bangla word breaks into several Token.
The practical results are three:
- The same thing written in Bangla uses more Token, so it costs more
- A Bangla answer takes a little longer to write
- Bangla text reaches the Context Window limit sooner
This is not a weakness of Bangla. It is the result of a decision made when the Models were built. Newer Models are better than before, but the difference is still there.
What Parameter means, without maths
Everyone hears "this Model has so many billion parameters", and nobody understands it. It can be understood without maths.
Imagine a huge mixing board with hundreds of millions of knobs. Each knob decides a little about which way the probability calculation leans for the next Token after seeing one Token. Training means turning those knobs, a little at a time, by showing text, until they sit in the right place. Parameters are those knobs.
Two things matter here. One, a single knob has no meaning by itself. The fact "where is Dhaka" is not written on any one particular knob. It is spread across the combined positions of hundreds of millions of knobs. So you cannot delete one fact from a Model, the way you can delete a row from Excel.
Two, more knobs does not mean more facts. More knobs means more fine detail and a better ability to catch complex relationships. A small Model built by looking at good text can run almost as well as a big Model on ordinary tasks.
| Small Model | Big Model | |
| Number of Parameters | Comparatively fewer | Many more |
| Speed | Fast | Slow |
| Cost | Low | High |
| Easy tasks, such as translation or summary | Almost equally good | Good |
| Multi-step reasoning or complex code | Weak | Clearly ahead |
| Running on your own laptop | Often possible | Usually not |
Training and inference: two different events
This is where the most misunderstanding happens. People think that whatever they tell ChatGPT, it learns. In reality there are two completely separate events.
Training means learning. It happened long ago, on the company's own servers, before you started using it. There, by showing a huge amount of text, the hundreds of millions of numbers inside the Model (these are called Parameters or Weights) were set in place. When the work is done, those numbers freeze and do not change.
Inference means use. You ask a question, and the answer comes out by calculating with those frozen numbers. Nothing inside the Model changes.
Here is a comparison: a student studied all year, then sat in the exam hall without books. Sitting in the hall he is not learning anything new. He writes with what had gone into his head. Your question is his exam question, not his year of study.
| Training (learning) | Inference (use) | |
| When it happens | Once, long before you use it | Every time you press Enter |
| Time it takes | Weeks to months | A few seconds |
| Where it happens | In the company's huge server farm | On a server, in response to your request |
| Cost | Huge, one time | Small, every time |
| Does what you write change the Model | No | No |
This is why a Model has a "knowledge cutoff", the last date of its information. It does not know by itself about events after that date. It cannot tell you today's market price or yesterday's match result, unless the app separately searches the internet and brings it in.
One more thing needs adding. It is also not true that it knows everything before the cutoff equally well. It knows well the subjects with a lot of writing on the internet, and less well the subjects with little. District-level or institution-level information about Bangladesh falls in this second group.
Is the internet stored inside the Model
No. This needs to be clear.
Inside the Model's file there is no saved text or page. There are only numbers, hundreds of millions of numbers, set by looking at text again and again. It is a bit like this: you have read newspapers for ten years, but the newspapers are not stored in your head. What you have is an idea of how each kind of news is written.
So when an AI states some fact, it is not "pulling it out" from somewhere, it is "making it". Most of the time the thing it makes is correct, because during learning it saw correct facts more often. But the door to error is always open.
Understanding this one point explains another: why it so often gives a wrong fact that is close to the right one. The book title is nearly right, but the author is someone else. The name of the law is right, but the section number is wrong. This is because it is not recalling the name, it is making something that looks like the name.
Context Window: how much room it has to remember
How does an AI in a chat remember what was said before?
The trick is very simple, and surprising to many people. Each time you write a new line, the app sends the whole earlier conversation to the Model again. So at the tenth question, the earlier nine questions and nine answers are all sent again. Each time, the Model reads everything with fresh eyes.
The amount of text it can hold in front of it at one time is called the Context Window, and it is measured in Token. Depending on the Model it can be from a few thousand to a million Token or more. For your own app, it is written in that app's information.
The window holds more than just what you wrote. What gets counted:
- The app's own hidden instructions, which you did not write but which take up room
- All the questions you have asked so far, more the longer the chat
- All the answers the AI has given so far, which are usually bigger than the questions
- Text extracted from uploaded files or images, and one PDF alone can eat a large part of the window
- The current question, the smallest part
- Room set aside for writing the answer, so when the window fills, the answer also gets shorter
Many people miss the last line. The window is not only room for reading, it is room for writing too.
A school administrator's long chat
An administrative officer at a school is working on an admission notice. The notice is twelve pages, about four thousand words, which by our assumed calculation is about 12,000 Token. He pastes the whole thing in.
Then in the same chat he asks 25 questions: how many seats in which class, when do forms start being given, what should I write for guardians, what will the short version for the notice board be. Each question and answer together is, say, 250 words, which is 750 Token.
The total comes to 12,000 plus 25 times 750, which is about 30,750 Token. For a Model with a big window this is nothing, and everything will run fine.
But what about a Model with a small window, or some limited free version, where the window is, say, eight thousand Token? The notice itself is bigger than the window. The app will quietly keep dropping the beginning. By the twentieth question, the first five pages of the notice are no longer in front of it. Yet it will not say "I don't know". It will answer confidently with whatever it can see.
A working rule comes out of this: if an answer suddenly does not match something said earlier, suspect the window first. The fix is to open a new chat and give only the part that is needed again, not all twelve pages.
Why it forgets
Now the reason for forgetting becomes clear. There are two different kinds of "forgetting".
One, forgetting in a long chat. As the talk grows and the Context Window fills, the early part starts being dropped. At the start you said "my shop's name is Sharighor", and fifty messages later it can no longer see that, because it has moved outside the window.
Two, forgetting in a new chat. A new chat is a completely empty window. Yesterday's conversation is not sent there at all, so there is no question of it knowing. Your words were not stored inside it, and never are.
Some apps now have a feature called "memory". It does not change the Model. It keeps some notes separately and quietly attaches them to your question every time. It feels as if it remembered, but the management is done from outside.
A practical tip: in a long piece of work, every now and then write the important facts again. "A reminder: the shop's name is Sharighor, delivery in Dhaka is 60 taka." This brings the thing back inside the window.
Memory, file upload and search: three separate arrangements
People mix these three up, but each does a different job. None of them changes the inside of the Model. All three are ways of putting extra text into the window.
| Arrangement | What it actually does | Does the Model change | Main limit |
| memory | Keeps some notes separately and quietly attaches them every time | No | You will not know what has been saved unless you check, and a wrong note can be saved too |
| Uploading a file or image | Pulls the text out of the file and puts it in the window | No | With a big file, not all of it goes in, only parts |
| Searching the internet | Brings a few pages and puts them in the window | No | The page may be wrong, or it may pick the wrong part |
| New chat | Empties the window completely | No | Loses all earlier context |
So "the AI has learned my file" is wrong. It has read the file, for that chat. When the chat ends, the reading ends.
Why the same question gets two different answers
When choosing the next Token, the Model does not just take the most likely one. It picks from the list of probabilities with a bit of randomness. A setting called Temperature controls how much of this randomness there is.
At a low setting the answer is dull but steady, and at a high setting it is creative but unsteady. This is why if you ask the same question twice, you may not get exactly the same answer. This is not an error, it is by design.
Being wrong with confidence
The most dangerous habit of an AI is that it states a wrong thing as neatly as it states a right thing. In English this is called hallucination.
By now the reason should be clear. Its job is to place the likely next word, not to verify the truth. For the question "who wrote this book", a likely answer has the shape of a person's name. Even if that name is not correct, the sentence looks perfect.
Two more things need adding here.
One, inside it there is no needle that tells apart "I know" from "I don't know". Everything it says is the result of the same process. So the words "I am sure" are not its self-assessment. They are also just a likely sentence.
Two, in practice it leans toward being helpful. When asked, giving an answer is its natural behaviour, and saying "I don't have that information" is less natural. So in the empty spaces it often puts in something that fits.
Remember, there is no relationship between how firm the language of an AI answer is and how accurate the facts are.
Where each mistake comes from
When they see a mistake, many people think "the AI is bad". In fact different mistakes have different causes, and different fixes.
| What you see | Likely cause | What to do |
| Does not know recent events | knowledge cutoff | Check whether search is on, or give the facts yourself |
| Forgot what was said earlier | Window is full | New chat, give the needed part again |
| Name or number nearly right, slightly wrong | It is making it, not recalling it | Ask for the source, or verify yourself |
| Gives another country's rule when asked about Bangladesh | Little written about us | Check against a government site |
| Sum of a long list is wrong | It is not a calculator | Ask for a formula, do the sum in Excel |
| Two answers to the same question | Temperature | If in doubt, assume it is guessing |
| Answer stops suddenly | The Token set aside for the answer ran out | Write "continue", or ask for it in parts |
Common mistakes and their fixes
Users also make some mistakes again and again. All of them have simple fixes.
- Mistake: doing all the work in one chat. Result: slow, expensive, and the opening instructions get lost. Fix: a new topic means a new chat
- Mistake: asking without giving context. Saying "write a post", getting a bad result and being disappointed. Fix: write four things: who you are, who it is for, how long, and what tone
- Mistake: making it add up numbers. Fix: have it build a formula, and let Excel do the adding
- Mistake: asking "are you sure" to feel reassured. It often agrees, or changes its own correct answer. Fix: look for certainty not in it, but in the source
- Mistake: seeing one part of an answer wrong and throwing away the whole thing. Fix: keep the structure, verify the facts
- Mistake: assuming it will remember the earlier mistake. In a new chat it does not remember any correction you made. Fix: keep the instructions you need again and again in a note, and paste them at the start
Step by step: how to run a long piece of work
Putting all of the above together gives a working format. The Khulna shopkeeper's job of 120 descriptions runs best this way.
- Split the work. Not all 120 at once, but twelve batches of ten
- Write one standing instruction and keep it in your phone's notes. The shop's name, the tone, the number of lines, what should come at the end
- Open a new chat and paste that instruction, then give the details of the first ten products
- Fix the first result carefully, then show it the fixed sample and say "the rest in this style"
- When the ten are done, close the chat, and in a new chat paste the instruction again
- In every batch, check the prices and sizes yourself, because it may make those up
- Whatever keeps going wrong, add it to the instruction. The note will get a little better with each batch
- At the end, read the whole thing yourself once, before publishing
The most valuable part of this format is step two. A good written instruction does the window's work for you.
When not to use it
Once you understand the inner working, this list stands up by itself. Do not use an LLM for these tasks, even though it writes well.
- For the final word. How much money, which date, which section of a law: it is not the final source for these
- For anything about today, unless the answer has a link to the source
- To check the sums of a long list, that is a job for a calculator and Excel
- To analyse anything that would do harm if leaked. NID, customers' numbers, passwords, none of them
- Where the blame for a mistake falls on someone else. Advice to a patient, a legal opinion, exam marks
- Where you cannot verify the result. Sending an important letter in a language you cannot read is risky
- To shift the responsibility for a decision. "The AI said so" is not a reason
What practical gain do you get from knowing this
Once you understand the inner working, the way you use it changes:
- When you ask for facts, ask for the source. For numbers, dates, laws or prices, verify separately
- Give the context yourself. It does not know your business, class or institution, and it will know only as much as you write
- Avoid long chats. Open a new chat for new work, which keeps the window clean
- Be careful with recent events. It does not know anything after the cutoff, and if it cannot search, it will guess
- Trust it more for language work, less for facts. Organising writing, translation, explaining things simply: it is truly good at these
What an AI actually is inside is explained from the ground up in the post on what AI is.
Questions and answers
Can an LLM think? Not in the way a person does. It does not run on experience, wishes or beliefs. But it would also be wrong to shrink it by saying "it only says the next word", because in getting that guess right it has picked up many structures of reasoning. For practical purposes, assume this: it understands structure well, and it does not confirm truth.
Does the Model learn by itself from what I write? Not inside the chat, and not at once. The numbers inside the Model do not change with your writing. However, many companies keep users' conversations and may later use them to build new Models. That is a separate matter, and it depends on your settings and plan.
Are Token and words the same thing? No. A Token is a piece of text. In English one word is roughly one Token on average, and in Bangla one word is often several Token. So "how many words" and "how many Token" are two separate counts.
What is the difference between a small Model and a big Model? Roughly, the number of Parameters and the amount of training. A big Model is usually better at complex work, but slower and more expensive. A small Model is fast and cheap, and enough for easy work. You do not need the biggest one for every job.
If an AI can search the internet, does it stop being wrong? It gets less wrong, but does not stop. The page it fetched may be wrong, or it may reach a wrong conclusion even from the right page. If there is a link to the source, it is safer to click it and look.
If the Context Window is big, are all the problems solved? Not all. A big window holds more text, that is true. But the fuller the window gets, the thinner its attention on any one particular line inside. It may skip a condition in the middle of a very long text. So even with a big window, it is a good habit to put the needed part in front separately.
If I ask "are you sure", does the truth come out? Not reliably. Often it politely agrees and changes its own correct answer, and often it states the wrong answer even more strongly. This question is not a self-check by it, it is just another question. Certainty has to be looked for in the source.
Is the answer better if I ask the same question in English? For hard or technical subjects, often yes, because the Model has seen far more English text. For easy subjects the difference is much smaller now. A useful trick is to give the instruction in English and ask for the answer in Bangla, which also uses fewer Token.
Can an AI catch its own mistakes? Somewhat, especially mistakes of reasoning or of the structure of a calculation. If you show it the answer and ask "tell me if there is any weakness here", you often get useful criticism. But it cannot catch a factual error by itself, because it has no source inside it to verify against.
In short
- An LLM, or Large Language Model, is a program that learned from a huge amount of text how likely each word is to follow another, and it writes answers by making that guess
- An AI breaks text into small pieces called Token; writing the same thing in Bangla takes more Token than English, so both cost and time go up
- Training means learning, which already happened and froze; inference means use, where nothing inside the Model changes, so your chat does not teach it
- No internet or database is stored inside the Model, only numbers; so it does not "fetch" facts, it "makes" them, and that is where the room for error is
- It forgets because of the Context Window: the whole conversation is sent again each time, the beginning drops off when the window fills, and a new chat means a completely empty window
- However firm the language of an answer is, it is not proof of accuracy; always verify dates, numbers, laws and prices separately