From one ordinary photo, AI usually cannot know accurately how many grams of food are present. An image shows shape, area, and some height, but it does not directly measure weight. To move from pixels to grams, a system must estimate volume and then infer the food’s density. Containers, utensils, and camera angle provide clues, but they cannot turn a camera into a food scale.
Portions shown in nosh come from visual inference and are useful for creating a quick, approximate log. If a decision requires exact grams, weighing the food or reading its package is more reliable.
A camera sees shape; a scale senses weight
A bowl of rice and a bowl of popcorn can occupy similar space while weighing very different amounts. A piece of watermelon and a piece of cake with the same visible size differ in water, air, and density. Seeing how large something looks does not directly tell you how heavy it is.
To estimate weight from a photo, AI generally has to infer three things in sequence: how much space the food occupies in the image, how large it may be in the real world, and how much that kind of food weighs per unit of volume. An error at any layer moves the final gram estimate.
A very specific-looking result such as “186 grams” may still be the output of several inferences. It does not mean the camera physically measured 186 grams.
A two-dimensional photo loses depth most easily
From directly above, you can see how widely rice spreads but not how high it is piled. From the side, height becomes clearer while food in front may block what is behind. Distance also changes apparent size: nearby food looks larger; food farther away looks smaller.
Deep bowls are especially difficult. A shallow layer of noodles and noodles filling the bowl to the bottom can look similar from above when the openings are the same size. Rice bowls, salads, and sandwiches stack ingredients too, leaving some food invisible in one image.
Recognizing more dish names cannot fully solve this. A two-dimensional image is missing part of the spatial information from the start.
Plates and utensils help, but they are not standard rulers
Keeping the full plate rim, takeout box, or cup in frame reduces some uncertainty about scale. Research systems sometimes use objects with known dimensions, coins, utensils, or containers to calibrate image-based portion estimates.
Everyday tableware does not have fixed dimensions. Two similar white bowls can have different diameters and depths; chopsticks, spoons, and hands vary by design and person. They offer rough proportion, not automatic exact grams.
Keeping the container is generally more useful than cropping the food to fill the screen. You still do not need to carry a ruler for every meal.
When a photo estimate is useful enough
If you want to know whether lunch contained a small or large bowl of rice, or compare the balance of staples and vegetables across several days, a photo estimate can remove much of the manual work.
After nosh returns a result, look first for a full-step portion error: half a bowl shown as a full bowl, one slice of bread as three, or a few bites of dessert as a complete serving. Those errors are worth correcting. Without measured weight, adjusting 140 grams to 155 merely makes the guess more detailed.
For everyday logs focused on direction, correct the errors that would change that direction.
When to use a scale instead
For packaged food, start with net weight and nutrition facts. When cooking and an accurate recipe matters, weigh the main ingredients. If you need strict intake control, are following a clinical nutrition plan, or use gram amounts to adjust medication or treatment, do not rely on photo estimates alone. Follow a doctor’s or qualified nutrition professional’s instructions.
You can edit the name, calories, and portion in a nosh result, but editing cannot replace a missing measurement. Use a known weight when you have one. When you do not, an approximate portion is more honest than an invented exact number.
A photo is good at showing what a meal roughly looked like. A food scale is good at answering how much it weighed. Let each tool do its own job instead of asking one image to provide every answer.