Paper: arXiv 2610.03406

Authors: Yuwei Wang, Hoi Ying Wong

Abstract

We propose an interactive robo-advising framework that learns personalized risk preferences from scores provided by clients. The resulting preference-learning problem is closely related to inverse reinforcement learning (IRL), as the robo-advisor infers the client’s latent reward specification from feedback. The robo-advisor interacts with clients iteratively as follows. At each interaction time, the advisor generates investment advice based on the optimal policy distribution derived from an inferred personalized risk preference. The client scores the advice. The advisor updates its assessment of the client’s risk preference based on the feedback. This learning procedure motivates us to investigate discrete-time Predictable Forward Exploratory Reward (PreFER) processes and derive an exploratory investment strategy. By interpreting the score as the acceptance probability of a piece of advice, our inverse learning procedure learns the client’s exploratory investment distribution using the acceptance-rejection method pioneered by von Neumann. Under CARA preferences, we show that, even though the scores contain noise, the robo-advisor can identify the client’s current risk aversion after a sufficiently large number of interactions. The PreFER process then carries the learned preference forward and generates future recommendations under updated market conditions.

Complexity vs Empirical Score

  • Math Complexity: 7.0/10
  • Empirical Rigor: 3.0/10
  • Quadrant: Lab Rats — theoretically deep, empirically untested

Why this score: The paper presents a mathematically sophisticated framework for learning client preferences in robo-advising, drawing connections to IRL and stochastic control. While theoretically strong, it lacks empirical validation or backtesting, focusing primarily on the theoretical derivation and identification under specific conditions. The novelty lies in the interactive scoring mechanism and the PreFER process.

Research Flowchart

  flowchart TD
    A[Research Goal: Personalized Robo-Advisor] --> B{Key Methodology: Interactive Learning};
    B --> C{Data/Input: Client Scores on Advice};
    C --> D[Computational Process: PreFER & Inverse RL];
    D --> E{Output: Updated Risk Preference & Exploratory Strategy};
    E --> F[Outcome: Identifiable Risk Aversion & Future Recommendations];