Files
Doomsday_Survival_Manual/active_skills/retry-on-fail/SKILL.md

176 lines
5.9 KiB
Markdown

---
name: retry-on-fail
description: Automatically stop executing any operation if the failure retry count exceeds 5 times. Use when Codex needs to perform repeated operations with automatic backoff and fail-safe behavior, such as network requests, file operations, or API calls that may temporarily fail.
---
# Retry on Fail
This skill provides guidance for implementing automatic retry logic with a maximum of 5 attempts before stopping execution.
## About the Skill
The `retry-on-fail` skill ensures robust operation by automatically retrying failed actions up to 5 times, then halting execution if all retries are exhausted. This prevents infinite loops and helps identify persistent failures that need manual intervention.
### When to Use This Skill
Use this skill when:
- Performing network requests (HTTP calls, DNS lookups)
- Executing file operations that may fail temporarily (write, read, delete)
- Making API calls with potential rate limits or timeouts
- Running commands in shell environments where transient errors occur
- Any operation where temporary failures are expected but persistent failures should stop execution
### Retry Behavior
- **Maximum attempts**: 5 retries total (including the initial attempt = 6 total tries)
- **Backoff strategy**: Exponential backoff with increasing delays between retries
- **Failure detection**: Based on error codes, HTTP status codes >= 400, or non-zero exit codes
- **Stopping condition**: After 5 failed attempts, execution stops and an error is reported
## Core Principles
### Concise is Key
The retry logic should be minimal and focused. Include only the essential code for:
1. Attempting the operation
2. Checking if it succeeded
3. Implementing backoff delay
4. Counting failures
5. Stopping after 5 attempts
### Set Appropriate Degrees of Freedom
Use a fixed maximum retry count (5) for consistency across all operations. This provides predictable behavior and makes debugging easier when persistent failures occur.
## Anatomy of a Skill
Every skill consists of a required SKILL.md file:
```
retry-on-fail/
└── SKILL.md (required)
├── YAML frontmatter metadata (required)
│ ├── name: retry-on-fail
│ └── description: [as above]
└── Markdown instructions (required)
```
## Skill Implementation Pattern
### Retry Logic Structure
The core retry mechanism follows this pattern:
1. **Initialize**: Set `attempt_count = 0`, `max_attempts = 5`
2. **Loop**: While `attempt_count < max_attempts`:
- Increment attempt counter
- Execute the operation
- If successful, return immediately
- If failed, apply backoff delay and continue
3. **Exhausted**: After all attempts fail, raise exception or return error
### Backoff Strategy
Use exponential backoff with the following delays:
- Attempt 1: Immediate (0s)
- Attempt 2: 1 second
- Attempt 3: 2 seconds
- Attempt 4: 4 seconds
- Attempt 5: 8 seconds
Total delay between attempts increases exponentially to avoid overwhelming the target system.
### Error Handling
- **Success**: Return immediately on first success
- **Transient failure**: Retry with backoff
- **Persistent failure**: After 5 attempts, stop and report error with full context
## Usage Examples
### Network Request Example
```python
def fetch_with_retry(url):
"""Fetch URL with retry logic"""
for attempt in range(6): # Initial + 5 retries
attempt_count += 1
try:
response = requests.get(url, timeout=10)
if response.status_code < 400:
return response.json()
except Exception as e:
pass
# Apply backoff (exponential)
delay = min(2 ** attempt, 8)
time.sleep(delay)
raise RetryExhausted(f"Failed after {attempt_count} attempts")
```
### File Operation Example
```python
def write_file_with_retry(filepath, content):
"""Write file with retry logic"""
for attempt in range(6):
attempt_count += 1
try:
with open(filepath, 'w') as f:
f.write(content)
return True
except Exception as e:
pass
# Apply backoff
delay = min(2 ** attempt, 8)
time.sleep(delay)
raise RetryExhausted(f"Failed to write file after {attempt_count} attempts")
```
### Shell Command Example
```python
def run_command_with_retry(cmd):
"""Run shell command with retry logic"""
for attempt in range(6):
attempt_count += 1
try:
result = subprocess.run(cmd, capture_output=True, timeout=30)
if result.returncode == 0:
return True
except Exception as e:
pass
# Apply backoff
delay = min(2 ** attempt, 8)
time.sleep(delay)
raise RetryExhausted(f"Command failed after {attempt_count} attempts")
```
## Best Practices
1. **Always track attempts**: Log the number of attempts made for debugging
2. **Use meaningful delays**: Exponential backoff prevents overwhelming targets
3. **Provide context on failure**: When stopping, include the last error and attempt count
4. **Consider retryable errors**: Only retry transient failures (network, timeout), not permanent ones (404, 500)
5. **Set appropriate timeouts**: Each individual operation should have its own timeout
## Common Pitfalls to Avoid
- **Infinite loops**: Always ensure a maximum attempt count is enforced
- **No backoff**: Linear delays can overwhelm systems; use exponential instead
- **Ignoring error types**: Not all errors are retryable (e.g., 404 vs. connection error)
- **Missing context**: Don't forget to log what failed and how many times
## Testing the Skill
Test the skill with:
1. **Transient failures**: Simulate network issues that resolve after retries
2. **Persistent failures**: Ensure it stops after 5 attempts
3. **Success on first try**: Verify it returns immediately without retrying
4. **Edge cases**: Test at exactly 5th attempt, varying delays