You can now describe an indicator in a sentence and watch it appear on a chart. That is a real advance in tooling, and it will produce a flood of indicators that look finished and are not. The problem is not the code. The code is usually fine. The problem is what the author of the sentence did not say, and what the machine therefore had to invent.
It cannot know what the data is
An indicator that "marks the previous day's high and low" needs to know which day. Exchange day, settlement day, or the calendar day of whatever time zone the platform happens to be in? For an index future those three answers put the line in three different places, and only one of them is the level the desk you are trading against is watching. A generated indicator picks one silently. It has to, because you did not tell it, and it cannot ask the tape.
The same goes for volume. "Volume" on a bar can be trades, contracts, or a feed's estimate from bid and ask changes. Delta reconstructed from bars is a guess wearing a number. An indicator written from a sentence inherits whichever definition the platform exposes first, and the sentence never mentioned it, so nobody checked.
It cannot know what it was for
A prompt says what to draw. It does not say what decision the drawing is supposed to support, so the generated logic optimises for looking correct on the chart it can see. That is how you get a "probability of revisiting the level" that is really a count of how often price crossed a line in the last sixty days, dressed up as a forecast. The number updates in real time, which makes it feel alive. It has no model of why the level would hold or break. It is a frequency, and a frequency from sixty days of one instrument is not a probability of anything.
It cannot know what it did not see
Anything written against bar data can only see bar data. Order flow, absorption, the sequence of trades inside a bar, the resting book: none of it exists in the inputs, so none of it can exist in the output, however confident the labels. An indicator that reports "absorption" from open, high, low, close and a volume total is naming a pattern in four numbers. Sometimes the pattern lines up with the real event. It has no way to know when it does not.
It cannot know whether it repaints
Generated code is written to satisfy the chart in front of it, which is a finished chart. Logic that uses the current bar's high or low before the bar has closed looks perfect on history and lies in real time. Nothing in the sentence "mark the swing highs" prevents that, and most generated versions do exactly it. The proof is a ten minute screenshot test, or a replay if your platform has one, and the people most likely to run the prompt are the least likely to run the test. We wrote up the test separately; it applies here without changes.
What a good tool does instead
- States its data definitions in the settings, so the session, the volume source and the time zone are decisions you made, not defaults you inherited.
- Reports what it measured, not what it inferred. A level is a level. A count is a count. A forecast is labelled as one and comes with the window it was drawn from.
- Never touches an unfinished bar for a value it will show as history.
- Keeps a diagnostics view that says, in plain words, the last thing that went wrong and where. Software that cannot report its own failure is asking you to trust it blind.
None of this is an argument against generated code. It is an argument against generated meaning. The sentence you type is the specification, and a specification that leaves out the data definition, the decision it serves, and the test that would falsify it is not a specification. It is a wish. The machine will grant it, and you will trade the result.