Mnesia:add_table_copy causes crashes when no replicas available

Vladislav Titov <[email protected]>
Newsgroups gmane.comp.lang.erlang.bugs
Message-ID <CA+stF05hev5Xd0+fkYU8S_hxPWwPXwV0_S9QjHBwGyXf-irtyQ@mail.gmail.com>
Hi

This is on a fully in-RAM mnesia cluster.

When performing mnesia:add_table_copy(tab, node(), ram_copies) when no
active replicas of tab are available, it replies back correctly with
{aborted, {system_limit, tab, ...}}.

However, looking at the mnesia_gvar afterwards, the {schema, local_tables}
key lists the table in it.

This (?) then causes mnesia to shut down when a node that is listed as
having the table gets started and adopts it as an orphan. In R14 (our
target release for now):

FATAL ** Sender failed: {error, {no_exists, tab}}

Or in R17:

FATAL ** Cannot load table foo from disc: {not_loaded, storage_unknown}

This then causes mnesia to shutdown.

The R14 case only happens if the node listed as actually holding the table
starts, and then re-starts. In R17 it seems to happen straight away on
first startup.

The active_replicas option for the table ends up listing the node() as well
at some point.

I've attached a repo case.

Any idea how I can work around this? Obviously checking active_replicas
before a copy would be a good idea, but it wouldn't protect against race
conditions when nodes are yo-yoing. Would cleaning up the local_tables gvar
be a good idea if system_limit happens?

thanks,
vlad

_______________________________________________
erlang-bugs mailing list
[email protected]
http://erlang.org/mailman/listinfo/erlang-bugs
bug_add_table_copy.escript (application/octet-stream, 1.8 KB) - not displayed
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.