Notes from Forja
Operational notes on AI in retail.
-
I know I need AI, I don't know where to start. Where do I start?
Start with the problem that already costs you a number, not with the technology. The best first project is the most expensive one you can measure today. If nothing is measured, that is project zero.
-
How Mercado Livre built a decade of AI in fraud and risk
It was not a project. It was ten years of narrow wins, each with a measured criterion. The small version does not buy the decade. It buys the first risk decision with a number.
-
The EU AI Act and the Brazilian retailer that exports: the reach nobody planned for
Europe's AI law has extraterritorial reach. If your AI system's output touches the European market, it reaches you, even with a Brazilian headquarters. The liability that enters through the back door.
-
Brazil's AI bill (PL 2338): what changes for those operating AI in retail
The bill classifies AI systems by risk. Credit scoring and biometrics fall in the high-risk band. Passed by the Senate, still in the Chamber. What to audit before it becomes law.
-
Internal data team or implementer: the real cost of each path
Build a team, hire someone who implements, or buy a managed service. The cost on the budget is the smallest of the three. The expensive part is time to first result and the dependency afterward.
-
How much of a retail AI project is estimate and how much is discovery
A large part of the first project's timeline is not estimate; it is discovery of the real data state. Whoever quotes a single number as if it were all estimable blows up when discovery arrives.
-
AI project ROI: the metric that looks healthy until the operation breaks
ROI is gain minus cost, over cost. The formula is trivial. The three lies live in the inputs: misattributed gain, the pilot's best case, and the maintenance cost nobody added.
-
What is a POC and why it does not prove what you think
A POC proves the technology can. It does not prove the operation will adopt. Confusing the two is where most AI projects die, between the pretty pilot and production.
-
Only 28% of AI projects deliver ROI. The number behind the number.
The number is broad enough to make a headline and narrow enough to hide what matters: of those that ship, most measured the criterion before the first line of code.
-
Internal team or implementer for the first AI project?
For the first project, an implementer, and use the project to form the internal team that takes over. Building a team before the first criterion is hiring for a target that does not exist yet.
-
The agent that worked in the demo and stalled in the operation
A retailer had a service agent that dazzled in the pitch and erred in real life. The problem was not the model. It was that the demo ran on a frozen catalog.
-
Walmart Jetblack: the text-message shopping concierge that did not close the math
Walmart built a shop-by-message service and closed it in 2020. The text looked like AI and was a lot of people. The lesson is not about conversation; it is about the math behind it.
-
AI customer service: build the agent, buy a bot platform, or implement
Three paths for customer service. The choice is not about features. It is about who answers the customer at 2am 18 months from now, and who answers when the bot gets it wrong.
-
How long until a WhatsApp agent stops promising what it does not have
With a narrow scope and real-time stock access, 8 to 14 weeks. The conversation is ready in days. What takes time is the wiring that stops the agent from lying.
-
WhatsApp as a sales channel: what the retail numbers show (and what they do not)
Magalu's Lu converts about 3x the app on WhatsApp. The number is real and misread. WhatsApp wins on reorder and quick questions, not long discovery.
-
In-store conversion rate: the number the store pretends to have
Conversion is buyers divided by visitors. A physical store almost never counts visitors, so it computes without the denominator and calls it conversion. E-commerce has it; the store fakes it.
-
Is scan-and-go worth it right now?
Worth it in high-traffic, large-basket formats where the line really hurts. But only after you have a loss criterion, because scan-and-go opens the same hole as self-checkout.
-
How Walmart made Sparky the discovery path in the app
Sparky does not sell because it chats well. It sells because it is wired into Walmart's catalog, stock, and delivery. The small version needs no Sparky. It needs the wiring.
-
What is an AI agent (and why it is not the same as a chatbot)
A chatbot answers. An agent acts: checks stock, opens an order, schedules. The difference is not better conversation; it is the action it is allowed to take. And the wrong action is expensive.
-
What is a dark store and when it stops making sense
A dark store is a store closed to the public, used only to assemble e-commerce orders. It is not a DC. And below a certain order density, picking from a live store is cheaper.
-
The category priced by habit. What changed when we measured.
A chain priced cleaning products by copying last year and following one competitor blind. Half the KVI list was wrong. The gain did not come from lowering price.
-
Personalized price and Brazil's CDC: the line between dynamic pricing and charging each customer differently
Changing the price for everyone at once is one thing. Charging each person differently is another, and it crosses the CDC and the LGPD. The difference defines your pricing project's risk.
-
Dynamic pricing in Brazilian retail: what the reports measure (and what they hide)
Every vendor publishes an adoption number. Most measure intent, and most are sponsored. Captured margin, the number that matters, rarely shows up.
-
Honest timeline: elasticity-based pricing in one category, 10 to 16 weeks
With a history that has real price variation, 10 to 16 weeks. With no past variation, there is no way to estimate elasticity, and the timeline goes open-ended. What doubles it.
-
Is a dynamic pricing engine worth it for a 30-store chain?
Worth it in one specific category first, usually perishables or general merchandise, before it pays as a chain-wide system. The gain lives where price is set by guess today.
-
Pricing: Pricefx, Revionics, build, or implement. Where each path wins.
Buy a pricing engine, build internally, or hire someone who implements. The choice is not about features. It is about who owns price on the next big promotion.
-
How Walmart prices millions of SKUs every day (and what to copy in a small chain)
Walmart's system is not auction pricing. It is low-price discipline with infrastructure to change price cheaply. The small version prices the KVI to competition and the tail to margin.
-
Markdown: the clearance that protects margin and the one that hides a buying error
Markdown measures how much margin you gave away in discount. Read as a total, it does not tell planned from forced discount. Same percentage, opposite meaning.
-
What is a KVI (known value item) and why it is more politics than data
A KVI is the item whose price the customer remembers and uses to judge whether the store is expensive. The list usually comes from the buyer's gut, not from sales. And the wrong list costs margin on both sides.
-
What is price elasticity and why lowering price does not always bring volume
Price elasticity is how much sales move when price moves. Measured as an average, it lies: it varies by SKU, store, week. Cutting price without it is margin that drops with no volume.
-
The expiry audit we got wrong in the first weeks
We built an expiry alert that flagged everything. The team ignored it. The mistake was not the model; it was measuring what the camera saw, not what the team acted on. Precision came later.
-
Dollar General pulled back on self-checkout. Shrink outran the savings.
The chain removed self-checkout from 300 stores and limited it in thousands. Self-checkout cut labor and opened loss. They never measured loss per store before scaling.
-
What a computer-vision shelf POC costs, and how long it takes
6 to 10 weeks to prove the concept on one section of one store, if ground truth exists. The POC that dies at scale is the one that skipped that condition. Where the timeline doubles.
-
Is computer vision on the shelf worth it for a 40-store chain?
Worth it for a specific problem, shelf-out on the fastest movers or expiry control, not as an all-seeing eye over the whole store. The camera is not the bottleneck. The operation is.
-
LGPD and store cameras: three things to audit before the next loss-prevention plan
An ordinary camera records personal data. Facial recognition records sensitive data, and Brazil's ANPD has already ordered one chain to shut it off. The gap between legal and enforced defines your risk.
-
How Amazon Go built Just Walk Out (and how much of it was people)
Amazon's invisible checkout was not 100% machine. The cost was never the camera; it was human review and the math that does not close on a grocery basket. The version that fits is different.
-
Service level reads one thing. The shelf does another.
Service level measures whether the DC delivered to the store. It stops at the back door. The missing pair is OSA, measured at the shelf, where the customer decides.
-
What is phantom inventory and why the system swears it has stock
Phantom inventory is when the system records units that do not exist in the store. It is silent shelf-out: replenishment never fires, because the system thinks it is stocked.
-
What is shrinkage and why three points vanish with no culprit
Shrinkage is the gap between the stock the system records and what actually exists. Read in aggregate, it hides where the leak is. And the leak has an address.
-
What is OSA (on-shelf availability) and why most chains measure it wrong
OSA is the percentage of items the customer finds on the shelf at the moment of purchase. Most chains measure it from system stock, and the system lies when phantom inventory exists.
-
How many months of history do you need for a reliable demand forecast?
For most categories, 18 to 24 months per SKU and store. Under 12 misses seasonality. But the amount matters less than how clean the data is.
-
The buyer trusted the spreadsheet over the system. He was partly right.
A regional chain asked us to make the buyer use the replenishment system. The buyer was ignoring a system that was learning wrong. The rebel was the load-bearing wall.
-
McKinsey projects $240 to 390B in AI for retail. How much of it is forecasting?
The report estimates potential value, not realized. The largest slice sits in replenishment, not the chatbot. And most projects never get there. The gap is the work.
-
Demand forecasting: platform, build, or implement. Where each path wins.
Buy RELEX or SAS, build internally, or hire someone who implements. The choice is not about features. It is about who operates the system on Black Friday 18 months from now.
-
How Lidl forecasts perishables by weather and calendar (and the version that fits your chain)
Lidl adjusts each store's fresh order against weather and calendar. The model is the cheap part. The 5% version runs in one pilot store.
-
Is it worth replacing the buyer's spreadsheet with automated forecasting?
For a 20 to 50 store chain, yes. But one category first, and without throwing away what the buyer's spreadsheet already knows. The spreadsheet is not the problem.
-
How long until SKU and store forecasting stops fighting the buyer
With a defined scope and data in reasonable shape, 14 to 22 weeks. Without that, however long it takes. What moves the timeline by 50% each way.
-
DSI reads one thing. The operation does another on holiday week.
DSI measures capital tied in stock, and measures it well. Read alone, it smooths the peak where shelf-out actually happens. The pair most dashboards miss.
-
What is stock cover and why it lies on perishables
Stock cover is how many days inventory lasts at today's sales rate. On perishables, it assumes the product never spoils. It does.
-
What is MAPE and where it misleads the retail operator
MAPE is the average percentage error of a demand forecast. The formula is honest. The aggregate reading hides shelf-out on the SKUs that pay the bills.
-
How Magalu's Lu became a WhatsApp salesperson that converts 3x the app
Lu converts better on WhatsApp than in Magalu's own app. But the channel and twenty years of brand move the number, not the model. Here's what you could copy.
-
Shelf-out has three causes. Only one is a forecasting problem.
We audited a regional chain's shelf-out SKU by SKU. Only a third was forecasting. The director was about to buy a better forecast to fix the wrong problem.
-
What is a forward-deployed engineer and why the term matters in retail
A forward-deployed engineer isn't a consultant or a remote developer. It's who writes code against your real data and hands the running operation back to your team.
-
McDonald's ended the IBM drive-thru AI. It lacked a number, not technology.
In 2021 McDonald's and IBM tested drive-thru voice ordering. In 2024 they ended it. The AI kept improving, but nobody set the number that lets the human leave.
-
Linear: issue tracking is dead. The 25% says more about retail AI than the 75%.
Linear: 75% of enterprise workspaces run agents; 25% of issues are now written by agents. The number behind it changes how retail AI projects need to be organized.
-
Forward-deployed engineer vs traditional consultancy vs AI platform
There are three paths to do AI in retail: platform, consultancy, or implementer. The right question is not which is best. It is which fits your operation.
-
Before the code, the criterion: what to measure in week one
Before the first line of code, three things have to be decided: the number, the range, and the review trigger. Without them, an AI project is aspiration, not a project.
-
Klarna didn't fail at AI. It failed to define when AI should have stopped.
In January 2024 Klarna announced $60M in savings from AI. In May 2025 it began rehiring. The AI worked. What was missing was the criterion for when it should have stopped.