takminの書きっぱなし備忘録 @はてなブログ

主にコンピュータビジョンなど技術について、たまに自分自身のことや思いついたことなど

2026/08/29第67回CV勉強会@関東「CVPR2026読み会」(後編)資料まとめ

8月29日に、前回に引き続き第67回コンピュータビジョン勉強会@関東「CVPR2026読み会」(後編)を、株式会社エクサウィザーズ様に会場をお借りして行いました。

以下、資料へのリンクまとめです。

登録サイト

kantocv.connpass.com

posfie

posfie.com

Youtube

www.youtube.com

コンピュータビジョン勉強会@関東

sites.google.com

資料まとめ

発表者 論文タイトル 発表資料
takmin Point Cloud as a Foreign Language for Multi-modal Large Language Model https://speakerdeck.com/takmin/point-cloud-as-a-foreign-language-for-multi-modal-large-language-model
herosan Efficiently Reconstructing Dynamic Scenes One D4RT at a Time https://drive.google.com/file/d/1R9G7DQr18DojdP_7--5-6ae8RYCgjyrU/view
こじま LagerNVS: Latent Geometry for Fully Neural Real-Time Novel View Synthesis
YutaKikuchi MAGICIAN: Efficient Long-Term Planning with Imagined Gaussians for Active Mapping https://speakerdeck.com/yutakikuchi_sd/cv-benkyoukai-kantou-magician-efficient-long-term-planning-with-imagined-gaussians-for-active-mapping
takubon INSID3: Training-Free In-Context Segmentation with DINOv3 https://drive.google.com/file/d/1HAALzH8QnLWiztboR6L7KwTvPFcogUoq/view
s-aiueo32 When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs https://speakerdeck.com/sansantech/260829
Hina39 IDperturb: Enhancing Variation in Synthetic Face Generation via Angular Perturbation https://speakerdeck.com/sansantech/260829-2
peisuke EgoX: Egocentric Video Generation from a Single Exocentric Video https://speakerdeck.com/peisuke/egox-egocentric-video-generation-from-a-single-exocentric-video
caprest Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems https://speakerdeck.com/caprest/dai-67-kai-konpyuta-bijon-benkyoukai-kantou-kouhen-scaling-aware-data-selection-for-end-to-end-autonomous-driving-systems
tomoaki_teshima HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps https://speakerdeck.com/tomoaki0705/holo-homography-guided-pose-estimator-network-for-fine-grained-visual-localization-on-sd-maps

takminの発表資料

speakerdeck.com

2026/07/18第67回CV勉強会@関東「CVPR2026読み会」(前編)資料まとめ

7月18日は、第67回コンピュータビジョン勉強会@関東「CVPR2026読み会」(前編)を、株式会社DeNA/IRIAM様に会場をお借りして行いました。

以下、自分で見返すために資料やリンク等をまとめておきます。

登録サイト

kantocv.connpass.com

posfie

posfie.com

Youtube

今回は途中で映像が流れなかったり、音声が流れなかったりしてしまい、申し訳ございません。

www.youtube.com

コンピュータビジョン勉強会@関東

sites.google.com

資料まとめ

発表者 論文タイトル 発表資料
alfredplpl Back to Basics: Let Denoising Generative Models Denoise https://www.docswell.com/s/alfredplpl/Z8N989-2026-07-18-082234
shunk031 PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow https://speakerdeck.com/shunk031/kantocv-67th-cvpr-2026
lychee1223_Lab SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers
Paper2Figure: A Multi-Agent Collaborative System for Figure Generation Towards Academic Research Paper
https://speakerdeck.com/lychee1223/kantocv-67th-cvpr-2026
kzykmyzw SoccerMaster: A Vision Foundation Model for Soccer Understanding https://speakerdeck.com/kzykmyzw/soccermaster-a-vision-foundation-model-for-soccer-understanding
Ito What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is Rarely F1 https://speakerdeck.com/keiichiito1978/20260618-cvprdu-mihui
welldone Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving https://speakerdeck.com/welldone/cvmian-qiang-hui-sensor2sensor
nasnetou Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers https://docs.google.com/presentation/d/13WFHhwXBmyhoa1bDJyeU8CvrrTLb-yZAacYtwr3eiw4/
Kenji AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models https://speakerdeck.com/tsukamotokenji/at-vla-b09cef8c-8208-4d1c-82e2-6cf35ec1ac84
TatsuyaSuzuki RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning https://speakerdeck.com/x_ttyszk/di-67hui-konpiyutabiziyonmian-qiang-hui-lun-wen-shao-jie-robowheel-a-data-engine-from-real-world-human-demonstrations-for-cross-embodiment-robotic-learning

後編は8/29(土)に(株)エクサウィザーズ様の会場をお借りして開催予定です。

kantocv.connpass.com

2026/02/08第66回CV勉強会@関東「世界モデル論文読み会」資料まとめ

2月8日は、第66回コンピュータビジョン勉強会@関東「世界モデル論文読み会」を、チューリング株式会社様に会場をお借りして行いました。

以下、自分で見返すために資料やリンク等をまとめておきます。

登録サイト

kantocv.connpass.com

posfie

posfie.com

Youtube

今回はネットワークが不安定だったため、動画が途切れ途切れになってしまい、申し訳ございません。

www.youtube.com

www.youtube.com

www.youtube.com

コンピュータビジョン勉強会@関東

sites.google.com

資料まとめ

発表者 論文タイトル 発表資料
tomoaki_teshima World Models (2018) / Dreamer (2019) Dropbox - kantocv_world_models.pdf - Simplify your life
Keiichi-Ito WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning 20260208_第66回 コンピュータビジョン勉強会 - Speaker Deck
Hidehisa Arai 世界モデルにおける分布外データ対応の方法論 世界モデルにおける分布外データ対応の方法論 - Speaker Deck
Godel 自律移動ロボットはWorld Modelsの夢は見るか?
大政孝充 GWM: Towards Scalable Gaussian World Models for Robotic Manipulation 第66回コンピュータビジョン勉強会@関東における株式会社ウェブファーマー 大政の発表 GWMモデル | PPTX
caprest 世界モデルで物を掴めるようになるのか? 世界モデルで物を掴めるようになるのか?.pdf - Google ドライブ
Kento Sasaki Epona: Autoregressive Diffusion World Model for Autonomous Driving (ICCV 2025) 第66回コンピュータビジョン勉強会@関東 Epona: Autoregressive Diffusion World Model for Autonomous Driving - Speaker Deck
takmin Cosmos World Foundation Model Platform for Physical AI Cosmos World Foundation Model Platform for Physical AI - Speaker Deck
Shin-kyoto Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail [CV勉強会@関東 世界モデル論文読み会] VLA自動運転model Alpamayo-R1 - Speaker Deck
abemii Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models [CV勉強会@関東 World Model 読み会] Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models (Mousakhan+, NeurIPS 2025) - Speaker Deck

takmin発表資料

speakerdeck.com

雪であまり人が集まらないかと心配しましたが、とても盛況で安心しました。

2025/11/16第65回CV勉強会@関東「ICCV2025読み会」資料まとめ

11月16日は、第65回コンピュータビジョン勉強会@関東「ICCV2025読み会」を、渋谷スクランブルスクエアの株式会社MIXI様に会場をお借りして行いました。

以下、自分で見返すために資料やリンク等をまとめておきます。

登録サイト

kantocv.connpass.com

Togetterあらためposfie

posfie.com

YouTube

www.youtube.com

コンピュータビジョン勉強会@関東

sites.google.com

資料まとめ

発表者 発表内容 資料
abemii World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model https://speakerdeck.com/abemii/cvmian-qiang-hui-at-guan-dong-cvpr2025-du-mihui-world4drive-end-to-end-autonomous-driving-via-intention-aware-physical-latent-world-model-zheng-plus-iccv-2025
caprest Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving https://docs.google.com/presentation/d/10JJiQdOwoKiw--6g7B5Mc2N2Y4BMpNKDCkQ5oSx5NZA
Shin End-to-End Driving with Online Trajectory Evaluation via BEV World Model https://speakerdeck.com/shinkyoto/cvmian-qiang-hui-at-guan-dong-iccv2025-wote-end-to-end-driving-with-online-trajectory-evaluation-via-bev-world-model
Keiichi-Ito AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction https://speakerdeck.com/keiichiito1978/20250916-di-65hui-konpiyutabiziyonmian-qiang-hui
alfredplpl ICCV 2025の動画生成を眺めてみた
StreamDiffusion: A Pipeline-level Solution forReal-Time Interactive Generation
VACE: All-in-OneVideo Creation and Editing
FiVE: A Fine-grainedVideo EditingBenchmark for EvaluatingEmerging Diffusion and Rectified Flow Models
https://www.docswell.com/s/alfredplpl/5VM392-2025-11-08-194529
ShuN Makino Is CLIP ideal? No. Can we fix it? Yes! https://speakerdeck.com/shun6211/lun-wen-shao-jie-is-clip-ideal-no-can-we-fix-it-yes-di-65hui-konpiyutabiziyonmian-qiang-hui-at-guan-dong
Kenji CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance https://speakerdeck.com/tsukamotokenji/di-65hui-konpiyutabiziyonmian-qiang-hui
peisuke Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos https://speakerdeck.com/peisuke/moto-latent-motion-token-as-the-bridging-language-for-learning-robot-manipulation-from-videos

今回、世界モデルに関する発表が多かったです。コンピュータビジョン業界の大きなトレンドになっているのかな?