Natural-language methods run from classical sentiment dictionaries through transformer embeddings to full LLM agents that read news and place trades. The signal side is well established (news tone, filing language, earnings-call transcripts); the agent side is new, fast-moving, and under-evaluated. Both sit here, alongside the retrieval and summarization work that supports research itself.

What to check when reading. The specific hazard of LLM papers is temporal leakage through the model: a model trained on data through 2024 “knows” what happened to a 2023 stock, so a backtest that prompts it about 2023 is contaminated. A credible paper controls for the training cutoff, fixes the model version, reports prompt and temperature, runs many seeds, and compares against a cheap text baseline. For sentiment-signal papers, the classic checks apply: point-in-time text timestamps, costs, and a horizon over which the signal is actually tradable.