Re: [PATCH 1/1] libceph: unpublish osdmap before destroying it on stop

Viacheslav Dubeyko <[email protected]> Mon, 20 Jul 2026 10:50:24 -0700
Newsgroups org.kernel.vger.ceph-devel
Message-ID <[email protected]>
On Mon, 2026-07-20 at 11:49 +0800, Ren Wei wrote:
> From: Zihan Xi <[email protected]>
> 
> osdmap_show() takes osdc->lock for reading and then walks
> osdc->osdmap while formatting the debugfs osdmap file. During client
> teardown, ceph_osdc_stop() closes the OSD state under the write side
> of
> that lock, drops the lock, and only then destroys osdc->osdmap while
> the
> pointer is still published.
> 
> A concurrent debugfs reader can therefore acquire the read side after
> the
> stop path has dropped it and traverse an osdmap that is being freed
> or has
> already been freed.
> 
> Take the osdmap pointer out of osdc under the write lock before
> dropping
> that lock, then destroy the saved pointer after unlocking. This makes
> the
> write lock wait for any current osdmap_show() reader and makes later
> readers observe a NULL osdmap instead of a freed object. Also make
> queued
> OSD map messages return after taking the lock if teardown has already
> unpublished the osdmap.
> 
> Fixes: 76aa844d5b2f ("ceph: debugfs")
> Cc: [email protected]
> Reported-by: Vega <[email protected]>
> Assisted-by: Codex:gpt-5.4
> Signed-off-by: Zihan Xi <[email protected]>
> Reviewed-by: Ren Wei <[email protected]>
> ---
>  net/ceph/osd_client.c | 12 ++++++++++--
>  1 file changed, 10 insertions(+), 2 deletions(-)
> 
> diff --git a/net/ceph/osd_client.c b/net/ceph/osd_client.c
> index 2ff00070c181..70467eede336 100644
> --- a/net/ceph/osd_client.c
> +++ b/net/ceph/osd_client.c
> @@ -4145,8 +4145,11 @@ void ceph_osdc_handle_map(struct
> ceph_osd_client *osdc, struct ceph_msg *msg)
>  	bool pauserd, pausewr;
>  	int err;
>  
> -	dout("%s have %u\n", __func__, osdc->osdmap->epoch);
>  	down_write(&osdc->lock);
> +	if (!osdc->osdmap)
> +		goto out_unlock;
> +
> +	dout("%s have %u\n", __func__, osdc->osdmap->epoch);
>  
>  	/* verify fsid */
>  	ceph_decode_need(&p, end, sizeof(fsid), bad);
> @@ -4238,6 +4241,7 @@ void ceph_osdc_handle_map(struct
> ceph_osd_client *osdc, struct ceph_msg *msg)
>  bad:
>  	pr_err("osdc handle_map corrupt msg\n");
>  	ceph_msg_dump(msg);
> +out_unlock:
>  	up_write(&osdc->lock);
>  }
>  
> @@ -5268,6 +5272,8 @@ int ceph_osdc_init(struct ceph_osd_client
> *osdc, struct ceph_client *client)
>  
>  void ceph_osdc_stop(struct ceph_osd_client *osdc)
>  {
> +	struct ceph_osdmap *osdmap;
> +
>  	destroy_workqueue(osdc->completion_wq);
>  	destroy_workqueue(osdc->notify_wq);
>  	cancel_delayed_work_sync(&osdc->timeout_work);
> @@ -5279,6 +5285,8 @@ void ceph_osdc_stop(struct ceph_osd_client
> *osdc)
>  						struct ceph_osd,
> o_node);
>  		close_osd(osd);
>  	}
> +	osdmap = osdc->osdmap;
> +	osdc->osdmap = NULL;
>  	up_write(&osdc->lock);
>  	WARN_ON(refcount_read(&osdc->homeless_osd.o_ref) != 1);
>  	osd_cleanup(&osdc->homeless_osd);
> @@ -5290,7 +5298,7 @@ void ceph_osdc_stop(struct ceph_osd_client
> *osdc)
>  	WARN_ON(atomic_read(&osdc->num_requests));
>  	WARN_ON(atomic_read(&osdc->num_homeless));
>  
> -	ceph_osdmap_destroy(osdc->osdmap);
> +	ceph_osdmap_destroy(osdmap);
>  	mempool_destroy(osdc->req_mempool);
>  	ceph_msgpool_destroy(&osdc->msgpool_op);
>  	ceph_msgpool_destroy(&osdc->msgpool_op_reply);

The fix looks really good.

Reviewed-by: Viacheslav Dubeyko <[email protected]>

Thanks,
Slava.