SAP CPI Failed Message Reprocessing: A Self-Service Guide for Integration Operations Teams

Operations teams can reprocess most failed SAP CPI messages without a developer. This takes three things: messages that can be searched by business values, a stored payload that can be pushed again, and permissions that decide who can resend what. Most failures come from temporary problems, such as a receiver outage or an expired certificate, and not from defects in the integration logic. A resend fixes them. The code does not need to change.
Why does a failed SAP CPI message still need a developer ticket?
A failed message needs a developer ticket when the people who notice the failure cannot find the message or cannot resend it. The standard SAP Cloud Integration monitor shows processing logs. It does not offer a manual restart for a failed message, and it does not search payload content.
Here is a common scenario. At 8:40 on a Tuesday, a sales order from the CRM is missing in S/4HANA. The customer service lead knows the customer number. They do not know the iFlow name, the message ID, or the tenant. They raise a ticket. The ticket waits for an integration developer who can read the message processing log (MPL). The order stays blocked for hours.
Most of these tickets do not need engineering work. Typical root causes include:
- The receiver system was down for a short time.
- A certificate or credential was renewed late.
- A downstream system rejected calls during a maintenance window.
- A network or platform problem lasted a few minutes.
In each case, the correct fix is to send the same message again. When only a developer can do that, a few minutes of downtime become a half-day delay for the business.
This guide describes a day-two workflow that support analysts, platform owners, application managers and key users can run themselves. It has four steps: find the message, choose retry or reprocess, apply guardrails, and read volume patterns.
Step 1: How do you find a failed SAP CPI message using business terms?
You find a failed message by business terms only if the iFlow records those terms at runtime. SAP Cloud Integration indexes message IDs, correlation IDs, application message IDs and custom header properties. It does not run a free-text search across payloads. An iFlow that never writes the order number to the log cannot be searched by order number.
Make business values searchable at design time
Add a script step early in the iFlow that writes key business values to the message processing log:
import com.sap.gateway.ip.core.customdev.util.Message
def Message processData(Message message) {
def messageLog = messageLogFactory.getMessageLog(message)
def body = message.getBody(java.lang.String)
def xml = new XmlSlurper().parseText(body)
def orderNo = xml.'**'.find { it.name() == 'SalesOrderNumber' }?.text()
def customerNo = xml.'**'.find { it.name() == 'CustomerNumber' }?.text()
if (messageLog) {
if (orderNo) messageLog.addCustomHeaderProperty('SalesOrder', orderNo)
if (customerNo) messageLog.addCustomHeaderProperty('Customer', customerNo)
}
// Application ID makes the message findable by a single business key
if (orderNo) message.setHeader('SAP_ApplicationID', orderNo)
return message
}
Keep the list short. Record only values that operators need to search, and do not write sensitive personal data to custom headers. Header values are visible to everyone with monitoring access.
Search filters an operations tool should support
When the values are recorded, a monitoring layer can combine these filters:
| Filter | Who uses it | When it helps |
|---|---|---|
| Business value (order, customer, invoice number) | Business users, key users | A colleague reports a missing document |
| Sender or receiver system | Platform owners | One system fails across many interfaces |
| Message ID or correlation ID | Support analysts | An alert or another team supplies an ID. A correlation ID follows one transaction across several iFlows |
| iFlow name | Integration support | The owning interface is known |
| Status | All roles | Show only what needs action now |
| Message type, application ID, time range | All roles | Narrow a large result set |
Combined filters turn a long log scroll into one query, for example "failed messages from the CRM in the last four hours" or "all messages for customer 104233 today." A downloadable error report lets you send the case to an application team outside integration without screen captures.
Keep retention in mind. SAP Cloud Integration keeps message processing logs for a limited period, and payload trace data expires much sooner. If your investigation window is longer than the platform retention, the monitoring layer must keep its own copy of the log data.
Step 2: Should you retry or reprocess a failed SAP CPI message?
Retry when the message content is correct and the failure was temporary. Reprocess when the message was received and transformed but could not be delivered. Do not resend when the failure comes from bad data or a mapping error. The correct action depends on where processing stopped.
How automatic retry works in SAP CPI
Automatic retry needs persistence. SAP Cloud Integration can only try a message again if the message was stored first. There are two common patterns:
- JMS queue. The sender side writes the message to a JMS queue. A second iFlow consumes the queue. If delivery fails, the message stays in the queue with status Retry, and the JMS sender adapter tries again. You configure the retry interval, exponential backoff and a maximum interval. You can also set a dead-letter handling option so that a message that keeps failing stops blocking the queue.
- Data Store with an exception subprocess. An exception subprocess writes the failed message to a Data Store. A scheduled iFlow reads the entries and tries delivery again.
Both patterns apply only to asynchronous processing. In a synchronous interface, such as a synchronous SOAP or OData call, the error goes straight back to the sender, and the sender must resend. JMS queues also count against tenant capacity limits, so queue design is a platform decision and not a per-interface habit.
This is what "CPI retry automation" means in practice. The retry logic is designed once in the iFlow and then reused. After that, a retry is a button for the operator, not a development task.
Manual reprocess needs a stored payload
A manual reprocess re-pushes the target payload, which is the transformed version that was ready for the receiver. That payload must exist somewhere. SAP Cloud Integration does not store payloads by default, and trace mode is only for short diagnostic sessions. So the iFlow or the monitoring layer must persist the payload at the right point in processing. Without this, "reprocess" means asking the sender system to send the original message again.
Decision table: retry, reprocess or escalate
| What you see | Likely cause | Action |
|---|---|---|
| Receiver timed out or refused the connection | Target system unavailable | Confirm the receiver is up, then reprocess the stored target payload |
| Temporary error, correct content, message in Retry status | Short platform or network problem | Let the automatic retry run, or trigger a retry |
| HTTP 401 or 403 from the receiver | Expired credential or certificate | Fix the credential first, then reprocess |
| HTTP 400, mapping error or validation failure | Content or logic problem | Do not resend. Download the error report and send it to the owning team |
| Duplicate-key error from the receiver | Document already exists | Do not resend. Check the target system first |
The last two rows matter most. Resending bad data only creates a second failure. Self-service reprocessing exists to handle the recoverable cases quickly, so that developers spend their time on the failures that need them.
Check for duplicates before you resend
A connection timeout does not prove that the receiver did not process the message. The receiver may have posted the document and then failed to send its response. Before you reprocess, confirm that the document does not exist in the target system. Where the risk of duplicates is high, design the interface to be idempotent. SAP Cloud Integration provides an Idempotent Process Call step that skips a message ID it has already processed. Many receivers also reject a repeated external reference number. Idempotency is the technical control that makes a resend button safe to hand to business users.
Step 3: What guardrails make self-service reprocessing safe?
Two controls make self-service safe: status-based action rules and role-based access. The platform team configures both once. After that, operators only see the actions that apply to the message and to their role.
Status-based action rules
SAP Cloud Integration assigns each message a status. The platform team decides which statuses allow which action:
| Status | Meaning | Typical rule |
|---|---|---|
| Completed | Processed successfully | No resend. This prevents a duplicate order in S/4HANA |
| Retry | Persisted and waiting for another attempt | Retry allowed |
| Failed | Processing ended with an error | Retry or reprocess allowed, depending on the error type |
| Escalated | An error was raised to the sender or through an error end event | Integration team only |
| Processing | Still running | No action |
| Cancelled, Discarded, Abandoned | Stopped or given up by the runtime | Review by the integration team before any resend |
Role-based access
Separate the permissions into three levels:
- Read. See that a message exists, its status and its error text.
- Download. Open the payload and error reports. Payloads often contain personal or commercial data, so this right is usually narrower than read access.
- Retry or reprocess. Trigger a resend.
A typical split gives a customer service lead read access for their own sender system, a support analyst read and download rights, and the operations team all three. Scope each role by sender or receiver system, not only by action. The exact split is a governance decision for each company, and the monitoring tool must enforce it.
Audit and compliance
Record every resend: who triggered it, when, for which message, and with what result. This audit trail answers questions from auditors and application owners. It also protects the operator. Where payloads contain personal data, apply the same retention and access rules as the source systems.
Step 4: How do message volume patterns prevent the next failure?
Volume patterns show the cause behind repeated failures, which a list of single failed messages cannot show. An aggregate view groups message counts by time period and by sender or receiver system. It uses the same filters as message search, so you can compare periods and systems without exporting data to a spreadsheet.
Watch for these four patterns:
- Load spikes. A sender that usually sends a steady stream sends a large batch at month end. If failures cluster in the same window, the receiver may be throttling or timing out under load. This is a capacity problem, not a retry problem.
- A quiet sender. Volume from one system drops to zero during business hours. Nothing failed in CPI, because nothing arrived. Only a volume view shows this gap, so set an alert for missing traffic as well as for failures.
- Repeat failure windows. The same sender and time window appear in the failed list every week. This usually points to a scheduled job, batch window or maintenance slot that the teams must agree on.
- Rising retry counts. Messages that finally complete after many retries still show a weak receiver. Track retries, not only failures.
Metrics to prove the change works
Measure these values before and after you introduce self-service reprocessing:
| Metric | What it shows |
|---|---|
| Mean time to recover (MTTR) for failed messages | Time from failure to successful delivery |
| Share of failures resolved without a developer | How much work moved to operations |
| Share of developer tickets that needed a code change | How much noise is left in the development queue |
| Repeat failures per interface per month | Whether root causes are being fixed |
| Messages resent in error (duplicates) | Whether the guardrails work |
Worked example: the missing Tuesday sales order
This is how the workflow resolves the missing order without a ticket.
- Search. The customer service lead enters the customer number. One message comes back: status Failed, sender CRM, receiver S/4HANA.
- Read. The error shows a connection failure to the receiver at 02:15. The lead's role allows read access, so they can see this much.
- Hand over. The lead's role does not include reprocess. They send the message ID to the support analyst on shift, not a ticket to the development queue.
- Check. The analyst confirms that S/4HANA is available and that the order does not already exist in S/4HANA.
- Reprocess. The analyst re-pushes the stored target payload. The status changes to Completed, and the order appears in S/4HANA.
- Check the pattern. The aggregate view shows six more CRM messages failed between 02:00 and 02:30, all with the same connection error. The analyst checks and reprocesses each one, then sends the time window to the S/4HANA Basis team.
No developer was involved, and no iFlow was opened. The integration team learns about the problem from a note about a maintenance window, not from a queue of tickets.
Problem, solution and where operations are heading
Problem. Integration teams handle recoverable failures that only need a resend. Business users wait, and developers lose time to work that does not use their skills.
Solution. Design for operations from the start. Record business keys as searchable headers. Persist payloads so they can be delivered again. Build asynchronous interfaces with JMS-based retry and idempotency. Give operators a monitoring layer with status rules, role-based access and volume analytics.
Vision. Next, the monitoring layer will classify failures by itself. It will match an error signature to a known cause, retry the safe cases automatically within set limits, and send only new or data-related failures to people. Manual reprocessing becomes the exception. Developer tickets are then about real defects and improvements.
How Tarento's iVolve closes the SAP CPI monitoring gap
Standard SAP CPI monitoring gives little payload detail and no restart option, which is why resend requests turn into developer tickets. iVolve, Tarento's SAP-certified integration framework, includes Integration Inspector, an enhanced monitoring layer. It adds full error reports, one-click restart, and search by sender, receiver, message type and application ID. With role-based controls and runtime alerting on top of the tenant, operations teams can find and resolve recoverable failures themselves, while the platform team decides who can resend what.
Frequently asked questions
Can you manually restart a failed message in SAP CPI? Not from the standard monitor. SAP Cloud Integration retries messages automatically only when they are persisted in a JMS queue or Data Store. A manual restart needs a stored payload and a monitoring layer that can push it again.
What is the difference between retry and reprocess in SAP CPI? A retry runs the message again when the content is correct and the failure was temporary. A reprocess re-pushes the already transformed target payload to a receiver that was unavailable.
Why does my failed SAP CPI message not go to Retry status? The message is probably synchronous or not persisted. Only asynchronous messages stored in a JMS queue or Data Store can move to Retry status. Synchronous errors go back to the sender.
How do I search SAP CPI messages by order number? Write the order number to the message processing log as a custom header property, or set it as the SAP_ApplicationID header, in a script step. After that, you can search for the value in monitoring.
How do I prevent duplicate documents when I resend a message? Confirm that the document does not exist in the target system first. Where possible, design the interface to be idempotent, for example with the Idempotent Process Call step or a unique external reference that the receiver checks.
Which failed messages should never be resent? Messages that failed because of mapping errors, validation errors, HTTP 400 responses or duplicate-key errors. Resending them only creates another failure. Send them to the owning team with the error report.
Ready to cut failed-message tickets from your integration queue? Talk to Tarento's SAP Integration Suite experts about Integration Inspector and iVolve, and get a monitoring setup built for the people who run your interfaces every day.


