Same sentence, "make the AI understand our company" — so why is one quote ten times another?
An industrial consumables trading company wanted an AI support assistant. The brief was one sentence: make it understand our products and our rules. Three quotes came back: NT$120,000, NT$450,000, NT$1,500,000. Nobody was lying. They were selling three different things: Vendor A sold prompt engineering, Vendor B sold retrieval plus prompting plus guardrails, Vendor C sold fine-tuning. All three were fighting the same enemy: hallucination.
Term 1: Prompt engineering
In one line: writing how the AI should behave into a fixed operating manual.
Everyday analogy: the note taped to the counter for a new part-timer — "If you are unsure, say you will check. Never calculate a discount yourself." You have not trained anyone. You have written down the rules.
In your business: the prompt says "Answer only questions about our products; for any pricing question, reply that a sales rep will follow up." Editing it takes minutes.
Cost meaning: the cheapest of the four, typically 8–24 hours. Check your vendor's work against Anthropic's prompt engineering documentation. The ceiling: a prompt will never tell the AI what is in your warehouse.
Term 2: Fine-tuning
In one line: retraining a general model on your own data so it behaves more like your company.
Everyday analogy: after the note, you send that part-timer through three months of in-house training. Their tone and judgment come out looking like your company's — but they still do not know how many boxes are left in the warehouse today.
In your business: the trading company has 30,000 support conversations in a highly consistent style, so fine-tuning can lock the output to it. With 200 conversations, the effect is unmeasurable.
Cost meaning: the money is not in the training run, it is in data preparation: cleaning and labeling 30,000 conversations runs 100–300 hours, which is where the six-figure quote comes from. OpenAI's own documentation tells you to exhaust prompting and retrieval first.
Common misconception: most people assume "let the AI read our data" means fine-tuning. That is retrieval's (RAG) job: fine-tuning changes how it speaks, not what it knows.
Term 3: Guardrails
In one line: checks before and after the AI speaks, blocking what it must not say.
Everyday analogy: dual authorization at a bank counter. However good the teller is, past a certain amount a manager still has to swipe. Not distrust — one mistake costs too much.
In your business: this company has to block three things: no prices, no delivery-date commitments, no unverified specifications. On the input side, "how much" routes to a human; on the output side, any figure gets held.
Cost meaning: 16–40 hours, but it decides whether you dare let the AI face customers. Without it you only get "AI drafts, human sends."
Term 4: Hallucination
In one line: the AI states something that does not exist, in a very confident tone.
Everyday analogy: a salesperson terrified of saying "I don't know." Unsure of a spec, they will not go check; they produce a number that sounds reasonable. They were trained to keep talking, not to be right.
In your business: a customer asks the temperature rating of a bearing, the model has no data and invents one. The customer orders, installs it, burns it out — and the screenshot has your company's support badge on it.
Cost meaning: the cost is not the wrong answer, it is that the wrong answer becomes your commitment. Lowering it takes four things at once: retrieval to anchor answers, prompting that requires an admission of ignorance, guardrails on high-risk replies, and a UI that shows sources.
How the four relate
The AI is a new support hire:
- Prompt = the work rules on their desk (fastest to change)
- Retrieval (RAG) = the continuously updated catalog (what they know)
- Fine-tuning = internal training (how they speak and judge)
- Guardrails = review and the forbidden-phrase list (what they cannot say)
- Hallucination = the symptom when the first four are not done
The order runs cheapest to most expensive: prompting → retrieval → guardrails → (only if genuinely needed) fine-tuning. Nine out of ten SMBs are done after the first three; the price gap is not dishonesty, it is skipped steps.
What this means for budget, timeline and risk
- Budget: prompting 8–24 hours, retrieval build 40–120 hours, guardrails 16–40 hours, fine-tuning data preparation 100–300 hours, plus model usage at NT$1,500–8,000/month.
- Timeline: the first three launch in 4–8 weeks; fine-tuning adds 6–12 weeks.
- Risk and maintenance: run "AI drafts, human sends" for three months, and name someone to keep the retrieval documents current, or in six months the AI will confidently quote a discontinued part number.
Five questions to ask your vendor
- "How much of this quote is fine-tuning, and what changes if we use retrieval instead?"
- "What does the AI say when it cannot find an answer? Can you demonstrate that?"
- "Will replies cite which document they came from, and can we click through?"
- "When the catalog changes, who updates the system, and how long until it takes effect?"
- "Which questions route straight to a human, and do we control that list?"
FAQ
Fine-tuning or RAG?
If the AI does not know your data, choose RAG. Only if it knows but does not sound like you should you fine-tune. RAG first: cheaper, updates instantly, traceable to source.
Can we launch without guardrails?
Yes, but only internally or as "AI drafts, human sends." Fully automated and customer-facing without guardrails leaves the risk to chance.
Can hallucination be eliminated?
No. You can only lower the rate and make it visible. Showing sources beats chasing a perfect model.
Once the prompt is written, is it done?
It needs maintenance. Model version changes and new product lines both force edits, so treat prompts like code.
We are small with little data — can we do this?
Little data is exactly why you start with prompting and retrieval. Twenty PDF catalogs are enough. On thin data, fine-tuning only makes the model invent more.
Next step
If you are holding an AI quote you cannot parse, send it over. Within 30 minutes we will mark which lines are prompting, retrieval and fine-tuning, and which can wait. No fee.
- Email: [email protected]
- Phone: 0916-224-047
- LINE: @ufv9089p