- Published on
How Machine Learning Can Identify MEV Attacks
Listen to the full article:
- Authors

- Name
- Jagadish V Gaikwad
Maximal extractable value, or MEV, is value captured by influencing which transactions enter a block and the order in which they execute. Machine learning can help identify suspicious ordering patterns, bot behavior, and abnormal transaction relationships, but it works best as part of a broader detection system rather than as an automatic proof of malicious intent.
A useful MEV detection model combines transaction data, block structure, decentralized exchange activity, gas behavior, and wallet relationships. It then searches for patterns such as a trade arriving immediately before and after another user’s swap, repeated profitable behavior across blocks, or coordinated activity between multiple addresses.
What MEV attacks look like
MEV is not always an attack. Arbitrage, for example, can restore price consistency between markets and may be a normal form of market activity. The security concern arises when transaction ordering lets one participant impose a predictable cost on another participant or manipulate a protocol’s state.
The most recognizable examples include:
- Sandwich attacks: A bot places one transaction before a victim’s trade and another after it. The first transaction moves the price, the victim receives a worse execution price, and the final transaction closes the position.
- Front-running: A bot observes a pending transaction and submits a competing transaction with a more favorable position or higher ordering priority.
- Back-running: A transaction is placed immediately after another transaction to exploit a state change, price movement, liquidation, or oracle update.
- Liquidation races: Bots compete to execute undercollateralized lending positions before other participants.
- Oracle manipulation: A trader temporarily moves a market price or liquidity pool state in a way that influences a protocol’s pricing decision.
MEV is commonly described in terms of transaction inclusion, exclusion, and ordering. The Springer research paper on MEV mitigation describes these mechanisms as including front-running, back-running, and sandwiching. Flashbots’ analysis of sandwich attacks also shows why detection is difficult: private order flow and changing bot strategies can make public-chain data incomplete.
Why machine learning is useful
Traditional MEV detection often starts with deterministic rules. A rule might flag a swap if the same address trades the same asset immediately before and after a victim transaction. This approach is understandable and easy to audit, but it can miss variations in timing, routing, contract usage, and multi-wallet coordination.
Machine learning can identify combinations of signals that are difficult to encode manually. Instead of asking only whether two trades bracket a swap, a model can consider:
- The size of each trade relative to pool liquidity.
- The change in the pool’s price between transactions.
- The gas price or priority fee used by each participant.
- The time between related transactions.
- Whether the same addresses repeat the pattern across many blocks.
- Whether the transactions use the same contracts, tokens, or routes.
- Whether the suspected bot’s final position is profitable after fees.
- Whether the victim’s execution was measurably worse than expected.
This makes machine learning especially valuable when attackers alter their behavior without changing the economic structure of the attack.
A model can also rank suspicious activity rather than issuing a simplistic yes-or-no decision. A high-risk score may send an event to an analyst, trigger additional tracing, or temporarily increase monitoring. That is safer than treating every unusual transaction as malicious.
The data needed for MEV detection
The quality of a detection model depends heavily on the data available to it. A basic transfer history is not enough because MEV is fundamentally about ordering, state changes, and economic relationships.
Transaction-level data
A detector can begin with standard transaction fields:
- Sender and recipient addresses.
- Block number and transaction index.
- Gas limit, gas used, and priority fee.
- Input data and method signature.
- Value transferred.
- Success or failure status.
- Contract addresses called during execution.
The transaction index is particularly important. It reveals the order in which successful transactions executed within a block. A suspected sandwich often appears as a three-part sequence, although more complex attacks can involve several addresses and contracts.
Execution traces
Transaction traces show internal calls that are not visible from the top-level transaction alone. They can reveal token transfers, pool interactions, flash-loan repayment, and calls routed through aggregators.
Trace data helps answer questions such as:
- Did the suspected attacker acquire an asset before the victim swap?
- Did the victim receive less of the desired asset because the pool state changed?
- Did the attacker sell or unwind the position afterward?
- Did several contracts act as intermediaries?
Without traces, a detector may incorrectly treat a router or aggregator as the attacker because it appears in the transaction path.
Decentralized exchange state
A model needs context about the market in which a transaction occurred. Useful features include pool reserves, liquidity, token prices, pool type, and the direction of the swap.
A large trade in a deep pool may have little price impact. The same trade in a shallow pool may move the price enough to create a sandwich opportunity. A detector that ignores liquidity can produce many false positives.
Address and graph relationships
MEV activity is often easier to understand as a graph than as a list. Addresses, contracts, pools, and transactions can be represented as connected entities.
Graph features may include:
- Repeated interaction between an address and a specific pool.
- Shared funding sources.
- Common deployment contracts.
- Repeated use of the same transaction builder or router.
- Transfers between suspected searcher addresses.
- Relationships between a searcher, a builder, and a beneficiary.
This does not prove that related addresses belong to the same operator. It does, however, help investigators identify coordinated patterns that are invisible when each wallet is analyzed independently.
How a machine-learning pipeline works
A practical detection system usually has several stages.
1. Collect and normalize blockchain events
The system gathers blocks, transactions, logs, traces, token transfers, and decentralized exchange events. It then normalizes differences between protocols and chains.
For example, one decentralized exchange may emit a standard swap event while another records a trade through a custom contract. A reliable pipeline must map both into comparable concepts without discarding protocol-specific information.
2. Reconstruct transaction relationships
The system groups transactions by block, pool, token pair, route, address, and execution time. It then searches for economically meaningful relationships.
For a possible sandwich, the pipeline may look for:
- A suspected attacker trade before a user swap.
- A user trade that changes the pool price.
- A suspected attacker trade after the user swap.
- A net asset movement that suggests the attacker profited.
The order alone is not enough. The detector must also test whether the first and final transactions form a coherent position and whether the user’s execution was adversely affected.
3. Engineer features
Feature engineering converts raw blockchain events into values a model can compare. Examples include:
- Transaction distance within a block.
- Difference in priority fees.
- Relative trade size.
- Price impact before and after a transaction.
- Estimated victim slippage.
- Attacker inventory before and after the sequence.
- Repeated activity by the same address.
- Profit after gas and protocol fees.
Estimated profit should be treated carefully. Token prices can change between the transaction and the time of analysis, and some assets may have thin or unreliable markets. A detector should preserve the assumptions behind every estimate.
4. Apply a suitable model
Different models answer different detection questions.
| Detection approach | Best use | Strength | Main limitation |
|---|---|---|---|
| Rules and heuristics | Known sandwich or liquidation patterns | Easy to explain and audit | Brittle when tactics change |
| Supervised classification | Labeled examples of confirmed attacks | Can learn complex combinations of signals | Requires reliable labels |
| Unsupervised anomaly detection | New or rare behavior | Does not require complete attack labels | Unusual activity is not automatically malicious |
| Sequence models | Ordered transaction behavior over time | Captures temporal patterns | More difficult to interpret and validate |
| Graph models | Coordinated wallets and contracts | Connects related entities | Depends on complete relationship data |
Research on machine-learning approaches to DeFi security identifies sequence models, Isolation Forest-style anomaly detection, supervised classifiers, and clustering as possible techniques for transaction and behavior analysis. These categories are useful design options, not guarantees that any particular model will detect every MEV strategy.
5. Investigate and score alerts
The final stage should produce an explanation alongside a risk score. An alert might state that three transactions in one block involved the same pool, that the first and last transactions belonged to a recurring address cluster, and that the sequence generated a positive estimated return while worsening the middle trade’s execution.
This explanation lets a reviewer distinguish a probable sandwich from ordinary arbitrage or a routing transaction.
Training data is the central challenge
Supervised machine learning requires labels. In MEV research, labels may come from manually reviewed transaction sequences, published datasets, known bot addresses, or high-confidence rules.
Each source has weaknesses:
- A rule-generated label may teach the model to repeat the rule rather than understand the attack.
- A known bot list may become outdated when operators change addresses.
- Manual labels are expensive and may reflect inconsistent judgments.
- Public data may omit private transactions or builder-level information.
- A profitable trade is not necessarily an exploit.
The Flashbots State of Sandwich Attacks analysis highlights the importance of data coverage. When order flow moves from public mempools into private channels, observers may no longer see the complete sequence that produced an outcome. A model trained only on public pending transactions can therefore undercount attacks or learn patterns that do not generalize.
A stronger training process combines positive, negative, and uncertain examples. Confirmed attacks should be separated from ambiguous sequences rather than forcing every observation into a binary label.
Detecting attacks before execution
Post-execution detection is easier because the system can inspect the final block, state changes, and token movements. Pre-execution detection is more difficult because the transaction sequence is incomplete.
A real-time system might analyze:
- Pending transaction calldata.
- Expected pool price impact.
- The sender’s historical behavior.
- Competing transactions targeting the same pool.
- Gas bidding patterns.
- Whether a proposed transaction would complete a profitable bracket around another pending trade.
However, mempool visibility varies by chain and transaction path. Private order flow can prevent a detector from seeing the relevant transactions before inclusion. That means machine learning can support early warnings, but it cannot observe events that are never exposed to the monitoring system.
Another option is to detect adversarial intent at the contract level. The paper Unveiling Attacks Ahead of Exploit studies machine-learning classifiers for identifying potentially adversarial decentralized-finance contracts rather than relying only on individual adversarial transactions. Its approach reported strong results on the authors’ dataset, but those results should not be treated as universal performance: they depend on the dataset, features, labels, and evaluation design.
False positives and adversarial adaptation
A good detector must separate harmful MEV from legitimate or ambiguous activity.
For example, a market-making strategy may buy before and sell after another transaction without targeting that user. An arbitrage transaction may follow a large swap because the swap created a genuine price difference. A router may appear in many suspicious sequences while merely executing instructions on behalf of users.
False positives can harm legitimate traders, delay transactions, or cause monitoring teams to ignore future alerts. False negatives are also serious because sophisticated searchers can split activity across wallets, vary gas settings, use private order flow, or route through contracts that hide simple patterns.
Attackers may also adapt to the model itself. Once the most important features become known, they can reduce trade sizes, insert decoy transactions, alter timing, or distribute profits. Effective systems therefore need periodic retraining, drift monitoring, randomized review rules, and human investigation for high-impact decisions.
How to evaluate an MEV detector
Accuracy alone is a poor evaluation metric when confirmed attacks are rare. A model that labels nearly everything as normal may achieve high accuracy while missing the events that matter.
More useful measures include:
- Precision: The proportion of alerts that are genuinely relevant.
- Recall: The proportion of known attacks the system finds.
- False-positive rate: How often normal activity is incorrectly flagged.
- Detection latency: How quickly the system raises an alert.
- Economic coverage: The estimated value or number of affected trades identified.
- Calibration: Whether a risk score corresponds reasonably to observed likelihood.
- Robustness: Whether performance holds across chains, protocols, market conditions, and new attacker behavior.
Evaluation should use time-based splits rather than randomly mixing historical transactions. Random splits can allow nearly identical behavior from the same bot to appear in both training and test data, producing an overly optimistic result.
A practical decision framework
Teams building or buying an MEV detector can use the following sequence:
- Define the threat precisely. Decide whether the goal is to find sandwiches, liquidation races, oracle manipulation, or broader anomalous ordering.
- Identify the observation boundary. Document whether the system sees public mempool data, finalized blocks, traces, private order flow, or only protocol-level events.
- Start with explainable rules. Establish a transparent baseline before adding complex models.
- Build labels conservatively. Separate confirmed, probable, ambiguous, and benign activity.
- Add machine learning where it improves coverage. Use it to rank alerts, discover clusters, or identify variations that rules miss.
- Validate across time and protocols. Test against new blocks, new pools, and behavior from addresses absent from training data.
- Preserve evidence. Store the transactions, traces, feature values, model version, and assumptions behind each alert.
- Pair automated scoring with review. High-impact decisions should not rely on an unexplained prediction.
Conclusion
Machine learning can identify MEV attacks by learning relationships among transaction order, pool state, gas behavior, address activity, and economic outcomes. Its strongest use is not replacing blockchain analysis but extending it: rules explain known patterns, graph and sequence models reveal coordination, and anomaly detection can surface behavior that does not match historical norms.
The main limitation is incomplete visibility. Private order flow, changing bot strategies, ambiguous labels, and legitimate arbitrage all make MEV classification uncertain. A dependable system should therefore present evidence and confidence, measure performance across changing conditions, and treat machine learning as one layer in a broader monitoring and investigation process.
You may also like
- Crypto Wallet Security: How AI Is Improving Fraud Detection
- How AI Helps SaaS Reduce Churn and Increase Retention
- How AI-Powered Automation is Reducing Operational Costs for Startups
- Best Web Hosting Companies with High Affiliate Payouts in 2025
- How to Migrate Your Website from Shared to Cloud Hosting: A Complete Step-by-Step Guide

