Skip to content

Allow to continue on partial shard failures - #1777

Open
fhaldi wants to merge 5 commits into
jertel:masterfrom
fhaldi:ignore-shard-failures
Open

fhaldi wants to merge 5 commits into
jertel:masterfrom
fhaldi:ignore-shard-failures

Conversation

@fhaldi

@fhaldi fhaldi commented Oct 1, 2026 •

Copy link
Copy Markdown

Description

Elastalert would raise an error if a shard is unavailable from the query, this means an old index out of the time range of the query is red, and recent onces are green and therefore query would work, it would simply skip.

This means that any degraded cluster might not be able to have rules to work.

This forces the query to be run even if some shards are unavailable.

Checklist

  • I have reviewed the contributing guidelines.
  • I have included unit tests for my changes or additions.
  • I have successfully run make test-docker with my changes.
  • I have manually tested all relevant modes of the change in this PR.
  • I have updated the documentation.
  • I have updated the changelog.

Questions or Comments

  • Not sure additional unit test is necessary as it is a change of a current feature
  • Documentation isn't necessary as it is backends I suppose?
  • I'll add the changelog tomorrow, not sure of what to write though

Let me know if you need additional information or I must change something.

@jertel

jertel commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Because this is a change in error handling behavior, this will require an opt-in flag. The new flag will require documentation. Yes, unit tests are mandatory for all changes, even to existing logic. Finally, add the changelog entry following the same pattern used by all other changelog entries.

@fhaldi

fhaldi commented Oct 2, 2026

Copy link
Copy Markdown
Author

Ty for the feedback!

I've added the requirements, the opt-in flag is allow_queries_on_degraded_indices. Maybe we want to reduce its name length but I'm not sure it can stay meaningful.

I also reduced log size when it lists shards because for example a cluster with 200 indices with 5 shards each gonna print 1000 shards in one line, maybe we want to handle it another way though?

Lmk if I have to change something else

@jertel

jertel commented Oct 4, 2026

Copy link
Copy Markdown
Owner

Your changes are looking good. Please add the new param to schema.yaml.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants