176 lines
5.9 KiB
Markdown
176 lines
5.9 KiB
Markdown
---
|
|
name: retry-on-fail
|
|
description: Automatically stop executing any operation if the failure retry count exceeds 5 times. Use when Codex needs to perform repeated operations with automatic backoff and fail-safe behavior, such as network requests, file operations, or API calls that may temporarily fail.
|
|
---
|
|
|
|
# Retry on Fail
|
|
|
|
This skill provides guidance for implementing automatic retry logic with a maximum of 5 attempts before stopping execution.
|
|
|
|
## About the Skill
|
|
|
|
The `retry-on-fail` skill ensures robust operation by automatically retrying failed actions up to 5 times, then halting execution if all retries are exhausted. This prevents infinite loops and helps identify persistent failures that need manual intervention.
|
|
|
|
### When to Use This Skill
|
|
|
|
Use this skill when:
|
|
- Performing network requests (HTTP calls, DNS lookups)
|
|
- Executing file operations that may fail temporarily (write, read, delete)
|
|
- Making API calls with potential rate limits or timeouts
|
|
- Running commands in shell environments where transient errors occur
|
|
- Any operation where temporary failures are expected but persistent failures should stop execution
|
|
|
|
### Retry Behavior
|
|
|
|
- **Maximum attempts**: 5 retries total (including the initial attempt = 6 total tries)
|
|
- **Backoff strategy**: Exponential backoff with increasing delays between retries
|
|
- **Failure detection**: Based on error codes, HTTP status codes >= 400, or non-zero exit codes
|
|
- **Stopping condition**: After 5 failed attempts, execution stops and an error is reported
|
|
|
|
## Core Principles
|
|
|
|
### Concise is Key
|
|
|
|
The retry logic should be minimal and focused. Include only the essential code for:
|
|
1. Attempting the operation
|
|
2. Checking if it succeeded
|
|
3. Implementing backoff delay
|
|
4. Counting failures
|
|
5. Stopping after 5 attempts
|
|
|
|
### Set Appropriate Degrees of Freedom
|
|
|
|
Use a fixed maximum retry count (5) for consistency across all operations. This provides predictable behavior and makes debugging easier when persistent failures occur.
|
|
|
|
## Anatomy of a Skill
|
|
|
|
Every skill consists of a required SKILL.md file:
|
|
|
|
```
|
|
retry-on-fail/
|
|
└── SKILL.md (required)
|
|
├── YAML frontmatter metadata (required)
|
|
│ ├── name: retry-on-fail
|
|
│ └── description: [as above]
|
|
└── Markdown instructions (required)
|
|
```
|
|
|
|
## Skill Implementation Pattern
|
|
|
|
### Retry Logic Structure
|
|
|
|
The core retry mechanism follows this pattern:
|
|
|
|
1. **Initialize**: Set `attempt_count = 0`, `max_attempts = 5`
|
|
2. **Loop**: While `attempt_count < max_attempts`:
|
|
- Increment attempt counter
|
|
- Execute the operation
|
|
- If successful, return immediately
|
|
- If failed, apply backoff delay and continue
|
|
3. **Exhausted**: After all attempts fail, raise exception or return error
|
|
|
|
### Backoff Strategy
|
|
|
|
Use exponential backoff with the following delays:
|
|
- Attempt 1: Immediate (0s)
|
|
- Attempt 2: 1 second
|
|
- Attempt 3: 2 seconds
|
|
- Attempt 4: 4 seconds
|
|
- Attempt 5: 8 seconds
|
|
|
|
Total delay between attempts increases exponentially to avoid overwhelming the target system.
|
|
|
|
### Error Handling
|
|
|
|
- **Success**: Return immediately on first success
|
|
- **Transient failure**: Retry with backoff
|
|
- **Persistent failure**: After 5 attempts, stop and report error with full context
|
|
|
|
## Usage Examples
|
|
|
|
### Network Request Example
|
|
|
|
```python
|
|
def fetch_with_retry(url):
|
|
"""Fetch URL with retry logic"""
|
|
for attempt in range(6): # Initial + 5 retries
|
|
attempt_count += 1
|
|
try:
|
|
response = requests.get(url, timeout=10)
|
|
if response.status_code < 400:
|
|
return response.json()
|
|
except Exception as e:
|
|
pass
|
|
|
|
# Apply backoff (exponential)
|
|
delay = min(2 ** attempt, 8)
|
|
time.sleep(delay)
|
|
|
|
raise RetryExhausted(f"Failed after {attempt_count} attempts")
|
|
```
|
|
|
|
### File Operation Example
|
|
|
|
```python
|
|
def write_file_with_retry(filepath, content):
|
|
"""Write file with retry logic"""
|
|
for attempt in range(6):
|
|
attempt_count += 1
|
|
try:
|
|
with open(filepath, 'w') as f:
|
|
f.write(content)
|
|
return True
|
|
except Exception as e:
|
|
pass
|
|
|
|
# Apply backoff
|
|
delay = min(2 ** attempt, 8)
|
|
time.sleep(delay)
|
|
|
|
raise RetryExhausted(f"Failed to write file after {attempt_count} attempts")
|
|
```
|
|
|
|
### Shell Command Example
|
|
|
|
```python
|
|
def run_command_with_retry(cmd):
|
|
"""Run shell command with retry logic"""
|
|
for attempt in range(6):
|
|
attempt_count += 1
|
|
try:
|
|
result = subprocess.run(cmd, capture_output=True, timeout=30)
|
|
if result.returncode == 0:
|
|
return True
|
|
except Exception as e:
|
|
pass
|
|
|
|
# Apply backoff
|
|
delay = min(2 ** attempt, 8)
|
|
time.sleep(delay)
|
|
|
|
raise RetryExhausted(f"Command failed after {attempt_count} attempts")
|
|
```
|
|
|
|
## Best Practices
|
|
|
|
1. **Always track attempts**: Log the number of attempts made for debugging
|
|
2. **Use meaningful delays**: Exponential backoff prevents overwhelming targets
|
|
3. **Provide context on failure**: When stopping, include the last error and attempt count
|
|
4. **Consider retryable errors**: Only retry transient failures (network, timeout), not permanent ones (404, 500)
|
|
5. **Set appropriate timeouts**: Each individual operation should have its own timeout
|
|
|
|
## Common Pitfalls to Avoid
|
|
|
|
- **Infinite loops**: Always ensure a maximum attempt count is enforced
|
|
- **No backoff**: Linear delays can overwhelm systems; use exponential instead
|
|
- **Ignoring error types**: Not all errors are retryable (e.g., 404 vs. connection error)
|
|
- **Missing context**: Don't forget to log what failed and how many times
|
|
|
|
## Testing the Skill
|
|
|
|
Test the skill with:
|
|
1. **Transient failures**: Simulate network issues that resolve after retries
|
|
2. **Persistent failures**: Ensure it stops after 5 attempts
|
|
3. **Success on first try**: Verify it returns immediately without retrying
|
|
4. **Edge cases**: Test at exactly 5th attempt, varying delays
|