The value of Cisco's Splunk lies in the output of its searches: Either a search alert triggers something, a search reports something or a dashboard constisting out of searchs shows something.
Therefore it is critical that searches are run correctly.
However due to some rare events, some searches arent executed correctly. For example each splunk user, including the system-user "splunk-system-user", has a memory limit for a search. If a search becomes very big, that limit can be reached and the search is aborted. Or due to a long running data modell search and a rolling restart, some searches may be stopped.
To get an alert (e.g. via e-mail) of all the Splunk searches which are aborted, not completed, canceled or failed, you may use the following search, which searches in the internal indexes _audit and _internal:
SPL query:
index=_audit savedsearch_name=* has_error_warn=* fully_completed_search=* info!=completed user=splunk-system-user app!=splunk_archiver info!=bad_request
| append [search index=_internal sourcetype=splunkd DispatchManager Search not executed: reason="*" | eval component="DispatchManager - Search not executed" ]
| table _time,component,reason,user,usage,quota,savedsearch_name app has_error_warn fully_completed_search info user event_count scan_count search_startup_time searched_buckets user search_id id
| rename component AS "component - Issue" | sort - _time | head 100
To find skipped seraches, use:
index=_internal splunk_server=*SplunkHostnames* host=*SH Hostnames* TERM(savedsearch) sourcetype=scheduler
| extract pairdelim="|;,", kvdelim="=:"
| search sourcetype=scheduler status=* NOT "status=success"
| table _time, host,search_type, user, app, savedsearch_name, reason,status,concurrency_limit,concurrency_category
| stats count by app,savedsearch_name, reason,status
| sort - count


No comments:
Post a Comment