The problem is sample size. The calculation assumes enough responses that the proportions mean something, and a business with thirty customers and eleven responses produces a number that swings wildly on individual answers. Watching it move month to month would be tracking noise, and the appearance of a metric is worse than no metric because it invites decisions.
The second problem is that the number tells you nothing actionable. Knowing your score fell does not tell you why, and the diagnosis has to come from the accompanying comment rather than the figure. Which means the useful part of the exercise is the follow up question, and you can ask that without the score.
What produces more at your scale is asking better questions of fewer people. What almost stopped you from buying. What nearly made you choose somebody else. What would have made this better. Those produce specific answers you can act on this week, from five customers, where a score from fifty would still require interpretation.
The one situation where it earns its place early is as a routing mechanism rather than a metric. Asking the question after a job, and using a high answer to prompt a review request while a low answer routes to a conversation with you, is genuinely useful. That is not measurement, it is triage, and it works at any volume.
Revisit it when you have enough customers that the number would be stable and enough of them that individual conversations are no longer feasible. At that point tracking a consistent measure over time starts to reveal trends, which is what it was designed for. Before that, talk to people.
Whatever you ask, ask it at a consistent moment in the relationship rather than whenever it occurs to you. Responses collected immediately after delivery and responses collected three months later measure different things, and mixing them produces a figure that cannot be compared against itself over time. Pick the point, apply it consistently, and the answers become a series rather than a collection. Whatever you use, keep the responses themselves rather than only the resulting figure, because the comments are what remain useful in a year while the number is not.