AI stock picker reviews: run a paper test before you pay

An AI stock picker review can tell you whether an app is easy to use. It cannot prove that the picks work.

That takes a record made before the outcome is known: dated picks, fixed rules, a fair benchmark, losing periods, and costs. Without those pieces, a review may be little more than a tour of the interface followed by a screenshot of a few winners.

You do not need to connect a brokerage account to check the basic claims. A paper test, run with hypothetical money, can expose stale recommendations, disappearing picks, vague sell rules, and a fee hurdle that the advertised returns never address.

Save the claim before testing it

Start by recording exactly what the provider promises. Save the pricing page, methodology, trial terms, refund policy, and any performance chart. Note the date.

Marketing language tends to blur several different products:

  • A stock screener filters a market using stated criteria.
  • A research assistant summarizes filings, news, or estimates.
  • A ranking model scores securities but may not say when to trade.
  • A signal service gives entries, exits, or portfolio weights.
  • An automated tool can place trades or rebalance an account.

A good summary is not a profitable signal, and a ranked list is not a portfolio. The test has to match the product's actual claim.

The Securities and Exchange Commission brought enforcement actions against two investment advisers in 2024 over statements about their use of AI. According to the SEC release, the firms agreed to settle charges involving false and misleading claims. The cases do not prove that every AI investing tool is unreliable. They do show why the word "AI" should not count as evidence by itself.

Build a record the app cannot rewrite

For every pick, record:

1. the ticker and company name; 2. the date and time the pick appeared; 3. the quoted price source; 4. the direction of the signal; 5. any target, stop, holding period, or exit rule; 6. the model score or explanation shown at that moment; 7. whether the result is live, simulated, or backtested.

A screenshot helps, but a spreadsheet is easier to analyze. Keep removed picks in the record. If a provider refreshes a "top ten" list every day, yesterday's names do not stop counting because the page changed.

Do not fill in missing sell rules after a stock moves. If the service only publishes buy ideas, choose a neutral rule before the test begins, such as holding each paper position for the same number of trading days. Label that as your test rule, not the provider's strategy.

This is where many casual reviews go wrong. They pick an attractive recommendation, wait for a gain, and write about the winner. A fair test records the whole list first.

Choose the benchmark before seeing results

A benchmark is the alternative the tool is supposed to beat. It should resemble the market and risk the picks actually use.

A list dominated by large U.S. growth companies should not be compared with cash. A small-company strategy should not claim victory merely because it beat a large-company index during a brief stretch when small stocks rallied. International, sector, leveraged, and crypto strategies need similarly relevant comparisons.

Write down the benchmark and the reason for choosing it before calculating performance. Use the same start and end times for both. Include reinvested distributions when the benchmark data makes a total-return series available.

S&P Dow Jones Indices publishes SPIVA scorecards comparing active funds with benchmarks over several periods. An AI app is not the same as an active mutual fund, but the research makes one point hard to dodge: outperforming an index needs to be measured across time, not declared from a short winning run.

Thirty days can catch obvious recordkeeping problems. It cannot establish a durable edge. A longer test that includes different market conditions says more, especially if the strategy makes only a few picks.

Count every pick, not every trade you liked

Use one weighting rule across the sample. Equal weighting is often the cleanest paper-test default because it prevents the reviewer from putting more hypothetical money behind the winners after seeing the result.

Also decide how to handle these cases before they occur:

  • a pick that cannot be bought at the recorded price;
  • a stock that is halted or unusually illiquid;
  • several signals for the same company;
  • a pick removed without an exit notice;
  • dividends, splits, mergers, and delistings;
  • a signal published outside regular market hours.

The point is not to design the perfect strategy. It is to prevent the rules from changing whenever the outcome becomes inconvenient.

Accuracy, or win rate, is not enough. A system can win on most trades and still lose money if its losses are much larger than its gains. Record average gain, average loss, cumulative return, and the largest peak-to-trough decline. If the app takes concentrated positions, show that too.

Subtract the bill

Reviews often quote the subscription price but never put it next to the account size. That hides how steep the hurdle can be.

Suppose a hypothetical tool costs $49 a month and its signals create another $10 a month in estimated trading costs. The annual total is $708. On a $25,000 paper portfolio, that equals 2.832% of the starting balance before taxes or any investment loss. If the benchmark assumption is 7%, the tool would need roughly 9.832% under this simple one-year comparison just to finish level with the benchmark. The figures are examples, not expected returns.

Use the AI Stock Picker ROI Calculator to compare a hypothetical tool return with a benchmark after the annual subscription cost. The AI Bot Fee Calculator adds estimated monthly trading costs. Neither calculator verifies a provider's results or predicts future performance.

Spreads and market impact can matter even when commissions are zero. Taxes can change the result in a taxable account, particularly when signals create frequent short-term trades. A paper test should state which costs it includes rather than quietly reporting a frictionless return.

Check what the performance chart leaves out

A provider's chart deserves a few direct questions:

  • Did the results begin before the strategy was available to customers?
  • Are they backtested, paper traded, or based on live customer orders?
  • Were the model and rules fixed throughout the period?
  • Are failed companies and old picks still in the data?
  • Do returns include subscription charges, spreads, and other trading costs?
  • Is the comparison price available after the signal reaches a user?
  • Does the chart show drawdowns and losing months?

Backtests can help describe how a set of rules behaved in historical data. They are not live results. A model can be tuned to the same history used to judge it, producing a clean chart that does not survive new data.

Live records have problems too. A provider may highlight one model while leaving retired versions out of the story. Ask whether the reported history covers every strategy sold under the product name or only the current survivor.

Investor.gov's fees glossary explains the ordinary but unavoidable point that fees reduce returns. The site's artificial intelligence and fraud alert also warns against relying on AI-generated information to make investment decisions and describes fraud pitches that use AI claims. Those warnings are especially relevant when a review repeats a provider's promotional numbers without checking the record behind them.

Reviews can reveal things a return test cannot

Performance is only part of the product. A careful review can still add useful evidence about:

  • whether sources and update times are visible;
  • how the app corrects errors;
  • whether scores change without an audit trail;
  • what data the provider collects;
  • whether the tool can trade, withdraw, or only read an account;
  • how cancellation and data deletion work.

Test customer support with a specific methodology question. "How does the model work?" may invite a polished answer. "At what time is a daily score fixed, and can I download every historical score including removed tickers?" is harder to dodge.

Read the privacy policy and permissions before linking an account. A research-only trial does not need authority to move money. If the tool uses an application programming interface, check whether withdrawal access can be disabled and how credentials can be revoked.

Our AI Investor Watch covers fees, custody, benchmarks, and common marketing claims. The AI stock picker app checklist is a shorter screen for comparing products before a trial.

What a useful verdict looks like

A credible review does not need to crown a winner. It can say that the app was good at organizing filings but that its performance record was too short to judge. It can find that the signals beat a chosen benchmark during the test while suffering a larger drawdown. It can also conclude that the data was too incomplete to calculate a fair result.

That last answer is not a failure. If picks disappear, timestamps are missing, or the rules change without a history, the lack of evidence is itself useful information.

Keep the conclusion as narrow as the test. A paper result does not show what every customer achieved, and one market period does not prove that a strategy will keep working. It does give a reader something better than star ratings: a record that can be checked.

Educational only. This article provides general information and a hypothetical testing method. It is not personalized financial, investment, tax, legal, cybersecurity, or trading advice, and it does not recommend any stock, benchmark, app, or strategy.

Sources